(n : Nat) โ @Eq Nat (@HAdd.hAdd Nat Nat Nat (@instHAdd Nat instAddNat) n 0) n
- Declsecond_probeDeclaration kindtheorem
(n : Nat) โ @Eq Nat (@HAdd.hAdd Nat Nat Nat (@instHAdd Nat instAddNat) n 0) n
(n : Nat) โ @Eq Nat (@HAdd.hAdd Nat Nat Nat (@instHAdd Nat instAddNat) n 0) n
(n : Nat) โ @Eq Nat (@HAdd.hAdd Nat Nat Nat (@instHAdd Nat instAddNat) n 0) n
Pajor's inequality, with no finiteness assumptions: the traces of ๐ on A are at most as many as the subsets of A shattered by ๐. For a finite trace family this is a descent on the number of traces that never consumes the ground set; an infinite trace family forces infinitely many shattered singletons, and both sides are โค.
Dvir, Filmus and Moran, A Sauer-Shelah-Perles Lemma for Lattices, credit the Boolean lattice case to Pajor (Sous-espaces โโโฟ des espaces de Banach, Travaux en Cours 16, Hermann, Paris, 1985) and to Aharoni and Holzman, unpublished. Their Theorem 1.2 is the lattice form, for finite lattices with nonvanishing Mรถbius function: a family shatters at least as many elements as it has members. Reading that for a family of traces on a ground set is the standard translation into the language of set families, and the statement here carries no finiteness hypothesis, which theirs does.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, ((fun x => A โฉ x) '' ๐).encard โค {B | B โ A โง Shatters ๐ B}.encardAt Y = Bool this specialises to Pajor's inequality.
Asserted byโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, ((fun x => A โฉ x) '' ๐).encard โค {B | B โ A โง Shatters ๐ B}.encardโ {ฮฑ : Type u_1} {A : Set ฮฑ} (N : โ) (๐ : Set (Set ฮฑ)),
((fun x => A โฉ x) '' ๐).Finite โ
((fun x => A โฉ x) '' ๐).ncard โค N โ ((fun x => A โฉ x) '' ๐).encard โค {B | B โ A โง Shatters ๐ B}.encardโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {x : ฮฑ},
(fun x => A โฉ x) '' {C | C โ ๐ โง x โ C} โช (fun x => A โฉ x) '' {C | C โ ๐ โง x โ C} = (fun x => A โฉ x) '' ๐โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {x : ฮฑ},
x โ A โ Disjoint ((fun x => A โฉ x) '' {C | C โ ๐ โง x โ C}) ((fun x => A โฉ x) '' {C | C โ ๐ โง x โ C})โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, Shatters {C | C โ ๐ โง x โ C} B โ x โ Bโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, Shatters {C | C โ ๐ โง x โ C} B โ x โ BA family whose members all avoid x shatters only sets avoiding x.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, (โ C โ ๐, x โ C) โ Shatters ๐ B โ x โ BInserting a fixed element is injective on the sets avoiding it: the element can be removed again, recovering the argument.
โ {ฮฑ : Type u_1} (x : ฮฑ), Set.InjOn (insert x) {B | x โ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
((fun x => A โฉ x) '' ๐).Nontrivial โ โ x โ A, (โ C โ ๐, x โ C) โง โ C โ ๐, x โ CA point of A witnessing that one trace fails to contain another is a point at which ๐ splits.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A C C' : Set ฮฑ},
C โ ๐ โ C' โ ๐ โ ยฌA โฉ C โ A โฉ C' โ โ x โ A, (โ D โ ๐, x โ D) โง โ D โ ๐, x โ DA family whose members all avoid x is disjoint from any family of sets containing x.
โ {ฮฑ : Type u_1} {x : ฮฑ} {๐ฎ ๐ฏ : Set (Set ฮฑ)}, (โ B โ ๐ฎ, x โ B) โ Disjoint ๐ฎ (insert x '' ๐ฏ)If B is shattered both by the members of ๐ avoiding x and by the members of ๐ containing x, then ๐ shatters insert x B. This is the exchange step in the proof of Pajor's inequality.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ},
Shatters {C | C โ ๐ โง x โ C} B โ Shatters {C | C โ ๐ โง x โ C} B โ Shatters ๐ (insert x B)โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, ((fun x => A โฉ x) '' ๐).Infinite โ {B | B โ A โง Shatters ๐ B}.InfiniteThe singleton of a point at which ๐ splits is shattered by ๐.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (fun x => {x}) '' splitPointsโ ๐ A โ {B | B โ A โง Shatters ๐ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {x : ฮฑ}, Shatters ๐ {x} โ (โ C โ ๐, x โ C) โง โ C โ ๐, x โ CShatters read at Set ฮฑ, as an introduction rule; the โฉ-shaped companion of Shatters.of_forall_le.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (โ โฆB : Set ฮฑโฆ, B โ A โ โ C โ ๐, A โฉ C = B) โ Shatters ๐ AShatters read at Set ฮฑ, where โ is โฉ and โค is โ: every subset of A is the intersection of A with a member of ๐. Definitionally the same statement, stated in the shape every use site wants, so that consumers obtain an โฉ-typed equation directly.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A B : Set ฮฑ}, Shatters ๐ A โ B โ A โ โ C โ ๐, A โฉ C = BA family splitting A at finitely many elements has finitely many traces on A.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (splitPointsโ ๐ A).Finite โ ((fun x => A โฉ x) '' ๐).FiniteA trace of ๐ on A is determined by its restriction to the elements at which ๐ splits: elsewhere, membership in the trace is decided by A alone.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, Set.InjOn (fun x => x โฉ splitPointsโ ๐ A) ((fun x => A โฉ x) '' ๐)A set family ๐ shatters a set A if all subsets of A can be obtained as the intersection of A with some element of the set family. We also say that A is traced by ๐.
{ฮฑ : Type u_1} โ [SemilatticeInf ฮฑ] โ Set ฮฑ โ ฮฑ โ Propโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ : Set ฮฑ} {A : ฮฑ}, Shatters ๐ A โ โ B โ ๐, A โค Bโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ โฌ : Set ฮฑ} {A : ฮฑ}, ๐ โ โฌ โ Shatters ๐ A โ Shatters โฌ Aโ {ฮฑ : Type u_1} [inst : BooleanAlgebra ฮฑ] {๐ : Set ฮฑ} {A : ฮฑ}, Shatters ๐ A โ Shatters ((fun x => xแถ) โปยน' ๐) Aโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ : Set ฮฑ} [inst_1 : OrderBot ฮฑ], Shatters ๐ โฅ โ ๐.NonemptyThe elements of A at which ๐ splits: some member of ๐ contains them and some does not. These are exactly the elements whose singleton ๐ shatters.
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ Set ฮฑ โ Set ฮฑโ {ฮฑ : Type u_1} {n d : โ} {๐ : Set (Set ฮฑ)}, HasVCDimLE d ๐ โ d โค n โ โ(vcGrowth n ๐) โค (Real.exp 1 / โd * โn) ^ dโ {ฮฑ : Type u_1} {n d : โ} {๐ : Set (Set ฮฑ)}, HasVCDimLE d ๐ โ d โค n โ โ(vcGrowth n ๐) โค (Real.exp 1 / โd * โn) ^ dโ (d m : โ), 0 < d โ d โค m โ โ i โ Finset.range (d + 1), โ(m.choose i) โค (Real.exp 1 / โd * โm) ^ d
Truncating a binomial expansion at degree d can only decrease it, for t โฅ 0.
โ {m d : โ} {t : โ}, 0 โค t โ d โค m โ โ i โ Finset.range (d + 1), โ(m.choose i) * t ^ i โค (1 + t) ^ mUndoing the weighting t ^ i costs a single factor (m / d) ^ d, since t = d / m and each i in range is at most d.
โ {m d : โ},
0 < โd โ
0 < โm โ
โd โค โm โ
โ i โ Finset.range (d + 1), โ(m.choose i) โค
(โm / โd) ^ d * โ i โ Finset.range (d + 1), โ(m.choose i) * (โd / โm) ^ iAt t = d / m, the binomial base is bounded by exp 1 raised to the degree.
โ {m d : โ}, 0 < โm โ (1 + โd / โm) ^ m โค Real.exp 1 ^ dThe Sauer-Shelah inequality, with the sum indexed by Finset.range (d + 1).
โ {ฮฑ : Type u_1} {n d : โ} {๐ : Set (Set ฮฑ)}, HasVCDimLE d ๐ โ vcGrowth n ๐ โค โ k โ Finset.range (d + 1), n.choose kโ {ฮฑ : Type u_1} {n d : โ} {๐ : Set (Set ฮฑ)}, HasVCDimLE d ๐ โ vcGrowth n ๐ โค โ k โ Finset.Iic d, n.choose kThe Sauer-Shelah inequality: a family of VC dimension at most d traces at most โ k โค d, n.choose k sets on any finite set of size at most n. The proof here derives it from Pajor's inequality.
Simon, A Guide to NIP Theories, states the same bound as his Lemma 6.4, for the growth function of a class of VC dimension at most k, and records in the notes to that chapter that the lemma is implicit in Vapnik and Chervonenkis (1971) and was rediscovered independently by Shelah (1972) and by Sauer (1972). Dvir, Filmus and Moran state it in the introduction to A Sauer-Shelah-Perles Lemma for Lattices and cite the same three sources. Perles is carried in the name they give the lemma; a publication of his is cited by neither source.
โ {ฮฑ : Type u_1} {n d : โ} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
HasVCDimLE d ๐ โ A.Finite โ A.ncard โค n โ ((fun x => A โฉ x) '' ๐).ncard โค โ k โ Finset.Iic d, n.choose kโ {ฮฑ : Type u_1} {A : Set ฮฑ},
A.Finite โ โ (d : โ), {B | B โ A โง B.ncard โค d}.ncard = โ k โ Finset.Iic d, A.ncard.choose kโ {ฮฑ : Type u_1} {A : Set ฮฑ}, A.Finite โ {B | B โ A โง B.ncard โค 0} = {โ
}The subsets of A of size at most d + 1 split into those of size at most d and those of size exactly d + 1.
โ {ฮฑ : Type u_1} (d : โ) (A : Set ฮฑ),
{B | B โ A โง B.ncard โค d + 1} = {B | B โ A โง B.ncard โค d} โช {B | B โ A โง B.ncard = d + 1}โ {ฮฑ : Type u_1} (d : โ) (A : Set ฮฑ), Disjoint {B | B โ A โง B.ncard โค d} {B | B โ A โง B.ncard = d + 1}The ncard transfer of encard_image_inter_le_encard_shatters, for A finite. Internal: the general statement is the encard one, which needs no hypothesis.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
A.Finite โ ((fun x => A โฉ x) '' ๐).ncard โค {B | B โ A โง Shatters ๐ B}.ncardPajor's inequality, with no finiteness assumptions: the traces of ๐ on A are at most as many as the subsets of A shattered by ๐. For a finite trace family this is a descent on the number of traces that never consumes the ground set; an infinite trace family forces infinitely many shattered singletons, and both sides are โค.
Dvir, Filmus and Moran, A Sauer-Shelah-Perles Lemma for Lattices, credit the Boolean lattice case to Pajor (Sous-espaces โโโฟ des espaces de Banach, Travaux en Cours 16, Hermann, Paris, 1985) and to Aharoni and Holzman, unpublished. Their Theorem 1.2 is the lattice form, for finite lattices with nonvanishing Mรถbius function: a family shatters at least as many elements as it has members. Reading that for a family of traces on a ground set is the standard translation into the language of set families, and the statement here carries no finiteness hypothesis, which theirs does.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, ((fun x => A โฉ x) '' ๐).encard โค {B | B โ A โง Shatters ๐ B}.encardโ {ฮฑ : Type u_1} {A : Set ฮฑ} (N : โ) (๐ : Set (Set ฮฑ)),
((fun x => A โฉ x) '' ๐).Finite โ
((fun x => A โฉ x) '' ๐).ncard โค N โ ((fun x => A โฉ x) '' ๐).encard โค {B | B โ A โง Shatters ๐ B}.encardโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {x : ฮฑ},
(fun x => A โฉ x) '' {C | C โ ๐ โง x โ C} โช (fun x => A โฉ x) '' {C | C โ ๐ โง x โ C} = (fun x => A โฉ x) '' ๐โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {x : ฮฑ},
x โ A โ Disjoint ((fun x => A โฉ x) '' {C | C โ ๐ โง x โ C}) ((fun x => A โฉ x) '' {C | C โ ๐ โง x โ C})โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, Shatters {C | C โ ๐ โง x โ C} B โ x โ Bโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, Shatters {C | C โ ๐ โง x โ C} B โ x โ BA family whose members all avoid x shatters only sets avoiding x.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, (โ C โ ๐, x โ C) โ Shatters ๐ B โ x โ BInserting a fixed element is injective on the sets avoiding it: the element can be removed again, recovering the argument.
โ {ฮฑ : Type u_1} (x : ฮฑ), Set.InjOn (insert x) {B | x โ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
((fun x => A โฉ x) '' ๐).Nontrivial โ โ x โ A, (โ C โ ๐, x โ C) โง โ C โ ๐, x โ CA point of A witnessing that one trace fails to contain another is a point at which ๐ splits.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A C C' : Set ฮฑ},
C โ ๐ โ C' โ ๐ โ ยฌA โฉ C โ A โฉ C' โ โ x โ A, (โ D โ ๐, x โ D) โง โ D โ ๐, x โ DA family whose members all avoid x is disjoint from any family of sets containing x.
โ {ฮฑ : Type u_1} {x : ฮฑ} {๐ฎ ๐ฏ : Set (Set ฮฑ)}, (โ B โ ๐ฎ, x โ B) โ Disjoint ๐ฎ (insert x '' ๐ฏ)If B is shattered both by the members of ๐ avoiding x and by the members of ๐ containing x, then ๐ shatters insert x B. This is the exchange step in the proof of Pajor's inequality.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ},
Shatters {C | C โ ๐ โง x โ C} B โ Shatters {C | C โ ๐ โง x โ C} B โ Shatters ๐ (insert x B)โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, ((fun x => A โฉ x) '' ๐).Infinite โ {B | B โ A โง Shatters ๐ B}.InfiniteThe singleton of a point at which ๐ splits is shattered by ๐.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (fun x => {x}) '' splitPointsโ ๐ A โ {B | B โ A โง Shatters ๐ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {x : ฮฑ}, Shatters ๐ {x} โ (โ C โ ๐, x โ C) โง โ C โ ๐, x โ CShatters read at Set ฮฑ, as an introduction rule; the โฉ-shaped companion of Shatters.of_forall_le.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (โ โฆB : Set ฮฑโฆ, B โ A โ โ C โ ๐, A โฉ C = B) โ Shatters ๐ AShatters read at Set ฮฑ, where โ is โฉ and โค is โ: every subset of A is the intersection of A with a member of ๐. Definitionally the same statement, stated in the shape every use site wants, so that consumers obtain an โฉ-typed equation directly.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A B : Set ฮฑ}, Shatters ๐ A โ B โ A โ โ C โ ๐, A โฉ C = BA family splitting A at finitely many elements has finitely many traces on A.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (splitPointsโ ๐ A).Finite โ ((fun x => A โฉ x) '' ๐).FiniteA trace of ๐ on A is determined by its restriction to the elements at which ๐ splits: elsewhere, membership in the trace is decided by A alone.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, Set.InjOn (fun x => x โฉ splitPointsโ ๐ A) ((fun x => A โฉ x) '' ๐)Any collection of subsets of a finite set is finite.
โ {ฮฑ : Type u_1} {A : Set ฮฑ}, A.Finite โ โ (p : Set ฮฑ โ Prop), {B | B โ A โง p B}.FiniteA family of VC dimension at most d shatters only sets of size at most d.
โ {ฮฑ : Type u_1} {d : โ} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ}, HasVCDimLE d ๐ โ Shatters ๐ B โ B.ncard โค dA family of VC dimension at most d shatters only finite sets.
โ {ฮฑ : Type u_1} {d : โ} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ}, HasVCDimLE d ๐ โ Shatters ๐ B โ B.FiniteHasVCDimLE d ๐
d โค n
A set family ๐ has VC dimension at most d if all the sets it shatters have size at most d.
{ฮฑ : Type u_1} โ โ โ Set (Set ฮฑ) โ PropA set family ๐ shatters a set A if all subsets of A can be obtained as the intersection of A with some element of the set family. We also say that A is traced by ๐.
{ฮฑ : Type u_1} โ [SemilatticeInf ฮฑ] โ Set ฮฑ โ ฮฑ โ Propโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ : Set ฮฑ} {A : ฮฑ}, Shatters ๐ A โ โ B โ ๐, A โค Bโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ โฌ : Set ฮฑ} {A : ฮฑ}, ๐ โ โฌ โ Shatters ๐ A โ Shatters โฌ Aโ {ฮฑ : Type u_1} [inst : BooleanAlgebra ฮฑ] {๐ : Set ฮฑ} {A : ฮฑ}, Shatters ๐ A โ Shatters ((fun x => xแถ) โปยน' ๐) Aโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A B : Set ฮฑ}, A โ B โ Shatters ๐ B โ Shatters ๐ Aโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ : Set ฮฑ} [inst_1 : OrderBot ฮฑ], Shatters ๐ โฅ โ ๐.NonemptyThe elements of A at which ๐ splits: some member of ๐ contains them and some does not. These are exactly the elements whose singleton ๐ shatters.
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ Set ฮฑ โ Set ฮฑThe growth of a set family is the maximum number of sets it cuts out from any set of size at most n.
{ฮฑ : Type u_1} โ โ โ Set (Set ฮฑ) โ โโ {ฮฑ : Type u_1} {n : โ} {๐ : Set (Set ฮฑ)} {d : โ},
vcGrowth n ๐ โค d โ โ โฆA : Set ฮฑโฆ, A.Finite โ A.ncard โค n โ ((fun x => A โฉ x) '' ๐).ncard โค dThe cardinal Pajor inequality for determined traces: the traces of ๐ on A that are determined by finitely many points inject into the shattered subsets, with no hypotheses. Unlike the โโ-valued encard_determined_le_encard_shatters, this bounds genuine cardinalities; the full trace family cannot replace the determined traces (not_mk_image_inter_le_mk_shatters).
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
Cardinal.mk โ{t | t โ (fun x => A โฉ x) '' ๐ โง โ F, F.Finite โง โ t' โ (fun x => A โฉ x) '' ๐, t' โฉ F = t โฉ F โ t' = t} โค
Cardinal.mk โ{B | B โ A โง Shatters ๐ B}The full trace family cannot replace the determined traces: the half-line cuts over the rationals trace continuum-many sets while shattering only countably many.
Asserted byโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
Cardinal.mk โ{t | t โ (fun x => A โฉ x) '' ๐ โง โ F, F.Finite โง โ t' โ (fun x => A โฉ x) '' ๐, t' โฉ F = t โฉ F โ t' = t} โค
Cardinal.mk โ{B | B โ A โง Shatters ๐ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, Cardinal.mk โ(diagโ ๐ A) โค Cardinal.mk โ{B | B โ A โง Shatters ๐ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
(diagโ ๐ A).Infinite โ
Cardinal.mk
โ{t | t โ (fun x => A โฉ x) '' ๐ โง โ F, F.Finite โง โ t' โ (fun x => A โฉ x) '' ๐, t' โฉ F = t โฉ F โ t' = t} โค
Cardinal.mk โ(diagโ ๐ A)โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A t F : Set ฮฑ},
(โ t' โ (fun x => A โฉ x) '' ๐, t' โฉ F = t โฉ F โ t' = t) โ
t โ (fun x => A โฉ x) '' ๐ โ โ t' โ (fun x => A โฉ x) '' ๐, t' โฉ (F โฉ diagโ ๐ A) = t โฉ (F โฉ diagโ ๐ A) โ t' = tโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (diagโ ๐ A).Finite โ ((fun x => A โฉ x) '' ๐).Finiteโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A t t' : Set ฮฑ},
t โ (fun x => A โฉ x) '' ๐ โ t' โ (fun x => A โฉ x) '' ๐ โ โ {y : ฮฑ}, y โ diagโ ๐ A โ (y โ t โ y โ t')The traces of ๐ on A that are determined by finitely many points are at most as many as the shattered subsets: the restriction of encard_image_inter_le_encard_shatters to the determined traces. Its cardinal-valued sharpening, which is genuinely stronger than the โโ statement, is mk_determined_le_mk_shatters.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
{t | t โ (fun x => A โฉ x) '' ๐ โง โ F, F.Finite โง โ t' โ (fun x => A โฉ x) '' ๐, t' โฉ F = t โฉ F โ t' = t}.encard โค
{B | B โ A โง Shatters ๐ B}.encardPajor's inequality, with no finiteness assumptions: the traces of ๐ on A are at most as many as the subsets of A shattered by ๐. For a finite trace family this is a descent on the number of traces that never consumes the ground set; an infinite trace family forces infinitely many shattered singletons, and both sides are โค.
Dvir, Filmus and Moran, A Sauer-Shelah-Perles Lemma for Lattices, credit the Boolean lattice case to Pajor (Sous-espaces โโโฟ des espaces de Banach, Travaux en Cours 16, Hermann, Paris, 1985) and to Aharoni and Holzman, unpublished. Their Theorem 1.2 is the lattice form, for finite lattices with nonvanishing Mรถbius function: a family shatters at least as many elements as it has members. Reading that for a family of traces on a ground set is the standard translation into the language of set families, and the statement here carries no finiteness hypothesis, which theirs does.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, ((fun x => A โฉ x) '' ๐).encard โค {B | B โ A โง Shatters ๐ B}.encardโ {ฮฑ : Type u_1} {A : Set ฮฑ} (N : โ) (๐ : Set (Set ฮฑ)),
((fun x => A โฉ x) '' ๐).Finite โ
((fun x => A โฉ x) '' ๐).ncard โค N โ ((fun x => A โฉ x) '' ๐).encard โค {B | B โ A โง Shatters ๐ B}.encardโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {x : ฮฑ},
(fun x => A โฉ x) '' {C | C โ ๐ โง x โ C} โช (fun x => A โฉ x) '' {C | C โ ๐ โง x โ C} = (fun x => A โฉ x) '' ๐โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {x : ฮฑ},
x โ A โ Disjoint ((fun x => A โฉ x) '' {C | C โ ๐ โง x โ C}) ((fun x => A โฉ x) '' {C | C โ ๐ โง x โ C})โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, Shatters {C | C โ ๐ โง x โ C} B โ x โ Bโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, Shatters {C | C โ ๐ โง x โ C} B โ x โ BA family whose members all avoid x shatters only sets avoiding x.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ}, (โ C โ ๐, x โ C) โ Shatters ๐ B โ x โ BInserting a fixed element is injective on the sets avoiding it: the element can be removed again, recovering the argument.
โ {ฮฑ : Type u_1} (x : ฮฑ), Set.InjOn (insert x) {B | x โ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
((fun x => A โฉ x) '' ๐).Nontrivial โ โ x โ A, (โ C โ ๐, x โ C) โง โ C โ ๐, x โ CA point of A witnessing that one trace fails to contain another is a point at which ๐ splits.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A C C' : Set ฮฑ},
C โ ๐ โ C' โ ๐ โ ยฌA โฉ C โ A โฉ C' โ โ x โ A, (โ D โ ๐, x โ D) โง โ D โ ๐, x โ DA family whose members all avoid x is disjoint from any family of sets containing x.
โ {ฮฑ : Type u_1} {x : ฮฑ} {๐ฎ ๐ฏ : Set (Set ฮฑ)}, (โ B โ ๐ฎ, x โ B) โ Disjoint ๐ฎ (insert x '' ๐ฏ)If B is shattered both by the members of ๐ avoiding x and by the members of ๐ containing x, then ๐ shatters insert x B. This is the exchange step in the proof of Pajor's inequality.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {B : Set ฮฑ} {x : ฮฑ},
Shatters {C | C โ ๐ โง x โ C} B โ Shatters {C | C โ ๐ โง x โ C} B โ Shatters ๐ (insert x B)โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, ((fun x => A โฉ x) '' ๐).Infinite โ {B | B โ A โง Shatters ๐ B}.InfiniteThe singleton of a point at which ๐ splits is shattered by ๐.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (fun x => {x}) '' splitPointsโ ๐ A โ {B | B โ A โง Shatters ๐ B}โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {x : ฮฑ}, Shatters ๐ {x} โ (โ C โ ๐, x โ C) โง โ C โ ๐, x โ CShatters read at Set ฮฑ, as an introduction rule; the โฉ-shaped companion of Shatters.of_forall_le.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (โ โฆB : Set ฮฑโฆ, B โ A โ โ C โ ๐, A โฉ C = B) โ Shatters ๐ AShatters read at Set ฮฑ, where โ is โฉ and โค is โ: every subset of A is the intersection of A with a member of ๐. Definitionally the same statement, stated in the shape every use site wants, so that consumers obtain an โฉ-typed equation directly.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A B : Set ฮฑ}, Shatters ๐ A โ B โ A โ โ C โ ๐, A โฉ C = BA family splitting A at finitely many elements has finitely many traces on A.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, (splitPointsโ ๐ A).Finite โ ((fun x => A โฉ x) '' ๐).FiniteA trace of ๐ on A is determined by its restriction to the elements at which ๐ splits: elsewhere, membership in the trace is decided by A alone.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ}, Set.InjOn (fun x => x โฉ splitPointsโ ๐ A) ((fun x => A โฉ x) '' ๐)A set family ๐ shatters a set A if all subsets of A can be obtained as the intersection of A with some element of the set family. We also say that A is traced by ๐.
{ฮฑ : Type u_1} โ [SemilatticeInf ฮฑ] โ Set ฮฑ โ ฮฑ โ Propโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ : Set ฮฑ} {A : ฮฑ}, Shatters ๐ A โ โ B โ ๐, A โค Bโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ โฌ : Set ฮฑ} {A : ฮฑ}, ๐ โ โฌ โ Shatters ๐ A โ Shatters โฌ Aโ {ฮฑ : Type u_1} [inst : BooleanAlgebra ฮฑ] {๐ : Set ฮฑ} {A : ฮฑ}, Shatters ๐ A โ Shatters ((fun x => xแถ) โปยน' ๐) AThe points of A on which ๐ is not constant.
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ Set ฮฑ โ Set ฮฑโ {ฮฑ : Type u_1} [inst : SemilatticeInf ฮฑ] {๐ : Set ฮฑ} [inst_1 : OrderBot ฮฑ], Shatters ๐ โฅ โ ๐.NonemptyThe elements of A at which ๐ splits: some member of ๐ contains them and some does not. These are exactly the elements whose singleton ๐ shatters.
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ Set ฮฑ โ Set ฮฑThe full trace family cannot replace the determined traces in mk_determined_le_mk_shatters: the half-line cuts over โ trace continuum-many sets while shattering only countably many.
ยฌCardinal.mk โ((fun x => Set.univ โฉ x) '' Set.range fun r => {q | โq < r}) โค
Cardinal.mk โ{B | B โ Set.univ โง Shatters (Set.range fun r => {q | โq < r}) B}The full trace family cannot replace the determined traces: the half-line cuts over the rationals trace continuum-many sets while shattering only countably many.
Asserted byยฌCardinal.mk โ((fun x => Set.univ โฉ x) '' Set.range fun r => {q | โq < r}) โค
Cardinal.mk โ{B | B โ Set.univ โง Shatters (Set.range fun r => {q | โq < r}) B}Function.Injective cutโ
{B | B โ Set.univ โง Shatters (Set.range cutโ) B}.Countableโ {B : Set โ}, Shatters (Set.range cutโ) B โ B.SubsingletonA set family ๐ shatters a set A if all subsets of A can be obtained as the intersection of A with some element of the set family. We also say that A is traced by ๐.
{ฮฑ : Type u_1} โ [SemilatticeInf ฮฑ] โ Set ฮฑ โ ฮฑ โ Propโ โ Set โ
Pajor's inequality for multiclass concept classes, with no hypotheses on the domain, the label type or the family: the traces of ๐ on S are at most as many as the pair-cubes of ๐ supported inside S.
The count runs on a split of the family at a point where two traces differ, taking one branch per realised value with no residual bucket. A cube of a branch never constrains the split point, so a cube pair-shattered by ฮผ branches is counted ฮผ times on the left and supplies 1 + ฮผ.choose 2 cubes on the right, which closes the accounting because ฮผ โค 1 + ฮผ.choose 2 for ฮผ โฅ 1, with equality exactly at ฮผ โ {1, 2}. An infinite trace family forces infinitely many one-point cubes and both sides are โค.
At Y = Bool this specialises to encard_image_inter_le_encard_shatters, which is encard_image_inter_le_encard_shatters_of_multiclass below. The pair-cube is Natarajan's shattering witness for spaces of functions, from p. 81 of On learning sets and functions; see the module docstring.
โ {X : Type u_1} {Y : Type u_2} (๐ : Set (X โ Y)) (S : Set X), (S.restrict '' ๐).encard โค (pairCubes ๐ S).encardAt Y = Bool this specialises to Pajor's inequality.
Asserted byBuilt on the same pair patterns and pair-shattering; neither implies the other.
Asserted byโ {X : Type u_1} {Y : Type u_2} (๐ : Set (X โ Y)) (S : Set X), (S.restrict '' ๐).encard โค (pairCubes ๐ S).encardInfinitely many traces force infinitely many cubes. Either the family disagrees at infinitely many points of S, each of which carries a one-point cube, or it realises infinitely many labels at one point, which carries infinitely many one-point cubes. The proof runs the contrapositive: finitely many cubes bound both the disagreement set and, above each of its points, the set of realised labels, and a trace is determined by its values on the disagreement set, so the traces embed in a finite product. This is what makes the multiclass inequality hypothesis-free, as infinite_setOf_shatters does in the binary case.
โ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} {S : Set X}, (S.restrict '' ๐).Infinite โ (pairCubes ๐ S).Infiniteโ {X : Type u_1} {Y : Type u_2} (z : X) (u : Y), Function.Injective (singCubeโ z u)โ {X : Type u_1} {Y : Type u_2} (z : X) (u v : Y), singCubeโ z u v z = pairOfโ u vโ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} {S : Set X} {u v : Y} {z : X},
z โ S โ (โ c โ ๐, c z = u) โ (โ c โ ๐, c z = v) โ singCubeโ z u v โ pairCubes ๐ Sโ {X : Type u_1} {Y : Type u_2} {u v : Y} (z : X), u โ v โ pairSupport (singCubeโ z u v) = {z}The whole accounting, as one injection from traces to cubes.
At a point where two traces differ, every realised value gets its own branch, and a trace is routed by the value it takes there. The branch injections supplied by the induction hypothesis land in cubes that avoid that point, and a cube pair-shattered by ฮผ branches carries ฮผ traces into ฮผ of the 1 + ฮผ.choose 2 cubes available above it: the cube itself for the branch selected as the anchor, and one graft for each of the other ฮผ - 1. Injectivity is read off the graft, since pairOf u ยท is injective and the pattern below the new point is recovered by evaluating away from it.
โ {X : Type u_1} {Y : Type u_2} (S : Set X) (N : โ) (๐ : Set (X โ Y)),
(S.restrict '' ๐).Finite โ
(S.restrict '' ๐).ncard โค N โ โ ฯ, Set.MapsTo ฯ (S.restrict '' ๐) (pairCubes ๐ S) โง Set.InjOn ฯ (S.restrict '' ๐)โ {Y : Type u_2} (u : Y), Function.Injective (pairOfโ u)Splitting off one realised value at a point of S where the family is not constant strictly lowers the number of traces, because a second value survives outside the branch.
โ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} {S : Set X} {x : X} {u : Y},
x โ S โ
โ {c : X โ Y},
c โ ๐ โ c x โ u โ (S.restrict '' ๐).Finite โ (S.restrict '' {d | d โ ๐ โง d x = u}).ncard < (S.restrict '' ๐).ncardTwo branches at one point of S cut disjoint families of traces.
โ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} {S : Set X} {x : X} {u v : Y},
x โ S โ u โ v โ Disjoint (S.restrict '' {d | d โ ๐ โง d x = u}) (S.restrict '' {d | d โ ๐ โง d x = v})โ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} {S : Set X} {P : X โ Set Y} {x : X} {u v : Y},
x โ S โ
P โ pairCubes {c | c โ ๐ โง c x = u} S โ
P โ pairCubes {c | c โ ๐ โง c x = v} S โ graft P x (pairOfโ u v) โ pairCubes ๐ Sโ {X : Type u_1} {Y : Type u_2} {S : Set X} {P : X โ Set Y} {x : X},
x โ S โ pairSupport P โ S โ โ (s : Set Y), pairSupport (graft P x s) โ SThe graft. A pattern pair-shattered by the branch at a and by the branch at b extends by one unconstrained point carrying the pair {a, b} there to a pattern pair-shattered by the whole family. The selection at the new point names the branch that realises it. This is the multiclass exchange step, and it mirrors Shatters.insert.
The two labels are not required to differ. Distinctness is what makes the result a pair pattern, and it enters at mem_pairCubes_graft_pairOf, not here.
โ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} {P : X โ Set Y} {x : X} {a b : Y},
P x = โ
โ
PairShatters {c | c โ ๐ โง c x = a} P โ PairShatters {c | c โ ๐ โง c x = b} P โ PairShatters ๐ (graft P x {a, b})โ {X : Type u_1} {Y : Type u_2} (P : X โ Set Y) (x : X) {s : Set Y},
s.Nonempty โ pairSupport (graft P x s) = insert x (pairSupport P)โ {Y : Type u_2} (u v : Y), pairOfโ u v = โ
โจ (pairOfโ u v).encard = 2โ {Y : Type u_2} {u v : Y}, u = v โ pairOfโ u v = โ
โ {Y : Type u_2} {u v : Y}, u โ v โ pairOfโ u v = {u, v}โ {X : Type u_1} {Y : Type u_2} {P : X โ Set Y},
IsPairPattern P โ โ (x : X) {s : Set Y}, s = โ
โจ s.encard = 2 โ IsPairPattern (graft P x s)โ {X : Type u_1} {Y : Type u_2} {P : X โ Set Y} {x : X}, P x = โ
โ graft P x โ
= Pโ {X : Type u_1} {Y : Type u_2} {๐ ๐ : Set (X โ Y)} {P : X โ Set Y}, ๐ โ ๐ โ PairShatters ๐ P โ PairShatters ๐ Pโ {X : Type u_1} {Y : Type u_2} {P P' : X โ Set Y} {x : X},
P x = โ
โ P' x = โ
โ โ {s s' : Set Y}, graft P x s = graft P' x s' โ P = P' โง s = s'โ {X : Type u_1} {Y : Type u_2} (P : X โ Set Y) (x : X) (s : Set Y), graft P x s x = sโ {X : Type u_1} {Y : Type u_2} {x z : X} (P : X โ Set Y) (s : Set Y), z โ x โ graft P x s z = P zThe single-point case: an admissible label at x extends to a selection taking it there.
โ {X : Type u_1} {Y : Type u_2} {P : X โ Set Y} {x : X} {u : Y},
u โ P x โ โ ฯ, ฯ x = u โง โ โฆz : Xโฆ, z โ pairSupport P โ ฯ z โ P zA choice of admissible labels prescribed on part of the domain extends to a selection for the whole pattern. The prescribed values arrive as a total function, so off the support there is always a label to copy and no hypothesis on Y is needed.
โ {X : Type u_1} {Y : Type u_2} {P : X โ Set Y} {B : Set X} {ฯ : X โ Y},
(โ โฆz : Xโฆ, z โ B โ z โ pairSupport P โ ฯ z โ P z) โ โ ฯ, (โ โฆz : Xโฆ, z โ pairSupport P โ ฯ z โ P z) โง Set.EqOn ฯ ฯ Bโ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} {S : Set X},
(S.restrict '' ๐).Subsingleton โ โ ฯ, Set.MapsTo ฯ (S.restrict '' ๐) (pairCubes ๐ S) โง Set.InjOn ฯ (S.restrict '' ๐)The unconstrained pattern is a cube of every nonempty family. It is the multiclass reading of shatters_bot.
โ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)} (S : Set X), ๐.Nonempty โ (fun x => โ
) โ pairCubes ๐ Sโ {X : Type u_1} {Y : Type u_2}, (pairSupport fun x => โ
) = โ
A pattern is a pair pattern when it offers either no constraint or exactly two labels at each point. The two labels are recorded as a set, so their order carries no information.
{X : Type u_1} โ {Y : Type u_2} โ (X โ Set Y) โ PropA family pair-shatters a pattern when every choice of one admissible label per constrained point is realised by some member.
{X : Type u_1} โ {Y : Type u_2} โ Set (X โ Y) โ (X โ Set Y) โ PropOverwrite a pattern at one point. Stated without a decidable equality on X, since the two branches are separated by a hypothesis rather than by a test.
{X : Type u_1} โ {Y : Type u_2} โ (X โ Set Y) โ X โ Set Y โ X โ Set YThe pair-cubes of ๐ over S: pair patterns supported inside S and pair-shattered by ๐.
{X : Type u_1} โ {Y : Type u_2} โ Set (X โ Y) โ Set X โ Set (X โ Set Y)The unordered pair {u, v}, degenerating to the empty set when u = v. Carrying the degenerate case inside the pair lets one formula cover both branches of the injection.
{Y : Type u_2} โ Y โ Y โ Set YThe points at which a pattern constrains a concept.
{X : Type u_1} โ {Y : Type u_2} โ (X โ Set Y) โ Set XThe cube supported at the single point z, offering the labels u and v there.
{X : Type u_1} โ {Y : Type u_2} โ X โ Y โ Y โ X โ Set YHasDSDimLE.hasNatarajanDimLE has no converse. The six-cycle on a two-point domain has Natarajan dimension 1, a Natarajan witness on both points being a four-cycle, while its whole trace family is a two-dimensional pseudo-cube, so its DS dimension is at least 2.
This is the hexagon Brukhim, Carmon, Dinur, Moran and Yehudayoff record after their Definition 6, and the four-cycle test the proof turns on is their Example 7. Their Theorem 2 pushes the separation to a class of Natarajan dimension 1 and infinite DS dimension; that construction rests on hyperbolic pseudo-manifolds and is cited in the module docstring rather than formalised here.
โ ๐, HasNatarajanDimLE 1 ๐ โง ยฌHasDSDimLE 1 ๐
Built on the same pair patterns and pair-shattering; neither implies the other.
Asserted byโ ๐, HasNatarajanDimLE 1 ๐ โง ยฌHasDSDimLE 1 ๐
HasNatarajanDimLE 1 sixCycleโ
โ {P : Fin 2 โ Set (Fin 6)},
PairShatters sixCycleโ P โ
pairSupport P = Set.univ โ โ {yโ yโ : Fin 6}, yโ โ P 0 โ yโ โ P 1 โ sixCycleEdgeโ yโ yโ = trueโ (a b c d : Fin 6),
sixCycleEdgeโ a c = true โ
sixCycleEdgeโ a d = true โ sixCycleEdgeโ b c = true โ sixCycleEdgeโ b d = true โ a = b โจ c = dDSShatters sixCycleโ Set.univ
sixCycleโ.Nonempty
โ (c : Fin 2 โ Fin 6),
sixCycleEdgeโ (c 0) (c 1) = true โ
โ (i : Fin 2), โ c', sixCycleEdgeโ (c' 0) (c' 1) = true โง c' i โ c i โง โ (j : Fin 2), j โ i โ c' j = c jA family whose members each have, in every direction, a neighbour in the family is DS-shattered on the whole domain.
โ {X : Type u_1} {Y : Type u_2} {๐ : Set (X โ Y)},
๐.Nonempty โ
(Set.univ.restrict '' ๐).Finite โ
(โ c โ ๐, โ (i : X), โ c' โ ๐, c' i โ c i โง โ (j : X), j โ i โ c' j = c j) โ DSShatters ๐ Set.univA family DS-shatters a set when its traces there contain a pseudo-cube.
{X : Type u_1} โ {Y : Type u_2} โ Set (X โ Y) โ Set X โ PropA family has DS dimension at most d when every set it DS-shatters has size at most d.
{X : Type u_1} โ {Y : Type u_2} โ โ โ Set (X โ Y) โ PropA family has Natarajan dimension at most d when every pair pattern it pair-shatters has support of size at most d.
{X : Type u_1} โ {Y : Type u_2} โ โ โ Set (X โ Y) โ PropA pattern is a pair pattern when it offers either no constraint or exactly two labels at each point. The two labels are recorded as a set, so their order carries no information.
{X : Type u_1} โ {Y : Type u_2} โ (X โ Set Y) โ PropA family of labellings of S is a pseudo-cube when it is nonempty and finite and every member has, in every direction, a neighbour in the family differing there and agreeing everywhere else. Over two labels the pseudo-cubes are exactly the Boolean cubes; over more labels there are others.
{X : Type u_1} โ {Y : Type u_2} โ {S : Set X} โ Set (โS โ Y) โ PropA family pair-shatters a pattern when every choice of one admissible label per constrained point is realised by some member.
{X : Type u_1} โ {Y : Type u_2} โ Set (X โ Y) โ (X โ Set Y) โ PropThe points at which a pattern constrains a concept.
{X : Type u_1} โ {Y : Type u_2} โ (X โ Set Y) โ Set XThe six-cycle read as a family of labellings of a two-point domain. Its traces form a two-dimensional pseudo-cube containing no two-dimensional Boolean cube, a six-cycle having no four-cycle.
Set (Fin 2 โ Fin 6)
The edges of a six-cycle on six labels.
Fin 6 โ Fin 6 โ Bool
Shelah's ded bound on the traces of a family of finite VC dimension. On an infinite ground set a family of finite VC dimension traces at most ded #A sets, and by mk_image_inter_range_cut_eq_ded the bound is attained. The infiniteness of A is not decorative: see exists_finite_ground_hasVCDimLE_not_mk_image_inter_le_ded.
โ {ฮฑ : Type u} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {d : โ},
HasVCDimLE d ๐ โ A.Infinite โ Cardinal.mk โ((fun x => A โฉ x) '' ๐) โค ded (Cardinal.mk โA)No cardinal function smaller than ded bounds the traces, which places ded at the exact strength of the bound.
Asserted byโ {ฮฑ : Type u} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {d : โ},
HasVCDimLE d ๐ โ A.Infinite โ Cardinal.mk โ((fun x => A โฉ x) '' ๐) โค ded (Cardinal.mk โA)Shelah's theorem. A family that traces more than ded #A sets on an infinite ground set A shatters finite subsets of A of every size. The proof takes a subset B โ A of least cardinality on which the family still traces more than ded #A sets, enumerates B by an initial ordinal so that every proper initial segment is smaller than B and therefore carries at most ded #A traces, prunes the tree of traces along that enumeration to the nodes at least (ded #A)โบ members pass through, and walks down the pruned tree collecting one point per step.
The conclusion cannot be strengthened to an infinite shattered set: the finite subsets of โ shatter every finite set and no infinite one.
The statement is due to Shelah. Both expositions used here, Adler, Introduction to theories without the independence property, Theorem 23, and Simon, A Guide to NIP Theories, Proposition 2.69, state it for the space of complete ฯ-types over a parameter set rather than for a set family, and the reading under which a complete ฯ-type over A is a trace on A and the independence property is the shattering of arbitrarily large finite sets is the standard translation of that statement, not a form either source asserts.
โ {ฮฑ : Type u} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ},
A.Infinite โ
ded (Cardinal.mk โA) < Cardinal.mk โ((fun x => A โฉ x) '' ๐) โ
โ (n : โ), โ S โ A, S.Finite โง S.ncard = n โง Shatters ๐ Sโ {ฮฝ : Cardinal.{u}} (w : ฮฝ.ord.ToType), Cardinal.mk โ(Set.Iio w) < ฮฝShelah's argument in the tree language: an enumeration of a ground set of size ฮผ all of whose initial segments carry at most ฮธ traces, with ded ฮผ โค ฮธ and more than ฮธ traces on the whole set, yields for each n a set of n indices on which ๐ cuts out every pattern. The ceiling ฮธโบ is regular, so the pruning of mk_sdiff_core_lt applies and the core keeps the whole width.
โ {ฮฑ : Type u} {๐ : Set (Set ฮฑ)} {W : Type u} {e : W โ ฮฑ} [inst : LinearOrder W] [WellFoundedLT W] {ฮผ ฮธ : Cardinal.{u}},
Cardinal.aleph0 โค ฮผ โ
Cardinal.mk W = ฮผ โ
Cardinal.aleph0 โค ฮธ โ
ded ฮผ โค ฮธ โ
(โ (w : W), Cardinal.mk โ((fun x => e '' Set.Iic w โฉ x) '' ๐) โค ฮธ) โ
ฮธ < Cardinal.mk โ((fun x => Set.range e โฉ x) '' ๐) โ
โ (n : โ), โ S, S.Finite โง S.ncard = n โง โ (ฯ : Set W), โ C โ ๐, โ x โ S, e x โ C โ x โ ฯโ (ฮบ : Cardinal.{u}), ฮบ โค ded ฮบThe level-w nodes of the tree of codes are the traces on the initial segment e '' Iic w.
โ {ฮฑ : Type u} {๐ : Set (Set ฮฑ)} {W : Type u} {e : W โ ฮฑ} [inst : LinearOrder W] (w : W),
Cardinal.mk โ(truncate w '' encโ e '' ๐) = Cardinal.mk โ((fun x => e '' Set.Iic w โฉ x) '' ๐)โ {ฮฑ W : Type u} {e : W โ ฮฑ} {C C' : Set ฮฑ} [inst : LinearOrder W] {w : W},
truncate w (encโ e C) = truncate w (encโ e C') โ โ v โค w, e v โ C โ e v โ C'Pruning to the core costs fewer than lam members: a member outside the core has a narrow truncation, the narrow nodes at one level are fewer than lam and carry fewer than lam members each, and there are fewer than lam levels. Regularity of lam closes both sums.
โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ]
{F : Set (Lex (W โ ฮฒ))} {lam : Cardinal.{u}},
lam.IsRegular โ
Cardinal.mk W < lam โ (โ (w : W), Cardinal.mk โ(truncate w '' F) < lam) โ Cardinal.mk โ(F \ coreโ lam F) < lamThe traces on the range of e are at most as many as the codes along e, two members with the same code agreeing at every index.
โ {ฮฑ : Type u} {๐ : Set (Set ฮฑ)} {W : Type u} {e : W โ ฮฑ},
Cardinal.mk โ((fun x => Set.range e โฉ x) '' ๐) โค Cardinal.mk โ(encโ e '' ๐)If g separates at most as much as f does on s, then f '' s is no larger than g '' s.
โ {ฮฒ ฮณ ฮด : Type u} {s : Set ฮฒ} {f : ฮฒ โ ฮณ} {g : ฮฒ โ ฮด},
(โ x โ s, โ y โ s, g x = g y โ f x = f y) โ Cardinal.mk โ(f '' s) โค Cardinal.mk โ(g '' s)โ {ฮฑ : Type u} {C C' D : Set ฮฑ}, D โฉ C = D โฉ C' โ โ a โ D, a โ C โ a โ C'โ {ฮฑ W : Type u} {e : W โ ฮฑ} {C C' : Set ฮฑ}, encโ e C = encโ e C' โ โ (w : W), e w โ C โ e w โ C'โ {ฮฑ W : Type u} {e : W โ ฮฑ} {C C' : Set ฮฑ} {w : W}, ofLex (encโ e C) w = ofLex (encโ e C') w โ (e w โ C โ e w โ C')The induction that produces the shattered points. In a two-valued tree whose core is wide, every wide node carries, for each n, a set of n indices above its own level on which the members of F below it realise every pattern. The step spends the wide node: lt_mk_big_fiber gives more than ฮผ wide nodes below it, exists_level concentrates them on one level, the pigeonhole on finite index sets returns two of them with a common tuple, and the index where the two disagree is the new point. Its position, strictly above the old level and at or below the new one, is what keeps the points distinct. This is the combinatorial form of the claim proved by induction in Adler, Theorem 23, and in Simon, Proposition 2.69.
โ {W : Type u} [inst : LinearOrder W] [WellFoundedLT W] {F : Set (Lex (W โ WithBot (ULift.{u, 0} Bool)))}
{ฮผ lam : Cardinal.{u}},
Cardinal.aleph0 โค ฮผ โ
Cardinal.mk W = ฮผ โ
lam.IsRegular โ
Cardinal.mk โ(F \ coreโ lam F) < lam โ
ded ฮผ < lam โ
(โ g โ F, โ (x : W), ofLex g x = valInโ โจ ofLex g x = valOutโ) โ
โ (n : โ) (w : W),
โ p โ bigAtโ lam F w,
โ S,
S.Finite โง
S.ncard = n โง
(โ x โ S, w < x) โง
โ (ฯ : Set W), โ g โ F, truncate w g = p โง โ x โ S, ofLex g x = valInโ โ x โ ฯA node inherits two-valuedness from the members it truncates.
โ {W : Type u} [inst : LinearOrder W] {F : Set (Lex (W โ WithBot (ULift.{u, 0} Bool)))}
{q : Lex (W โ WithBot (ULift.{u, 0} Bool))} {x v : W},
(โ g โ F, โ (y : W), ofLex g y = valInโ โจ ofLex g y = valOutโ) โ
q โ truncate v '' F โ x โค v โ ofLex q x = valInโ โจ ofLex q x = valOutโBelow the level of truncation the sequence is unchanged.
โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ] {v : W}
{g : Lex (W โ ฮฒ)} {x : W}, x โค v โ ofLex (truncate v g) x = ofLex g xThe finite subsets of an infinite W of a prescribed size are at most #W many. This is the pigeonhole that forces two nodes at one level to carry the same tuple.
โ {W : Type u}, Cardinal.aleph0 โค Cardinal.mk W โ โ (n : โ), Cardinal.mk โ{S | S.Finite โง S.ncard = n} โค Cardinal.mk WThe level found by exists_level lies strictly above the node's own level: at or below it the fibre is a single node, hence too small.
โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ]
{F : Set (Lex (W โ ฮฒ))} {lam ฮผ : Cardinal.{u}} {v w : W} {p : Lex (W โ ฮฒ)},
Cardinal.aleph0 โค ฮผ โ ฮผ < Cardinal.mk โ{q | q โ bigAtโ lam F v โง truncate w q = p} โ w < vThe step at which ded enters the argument. Below a wide node p, the core members extending p are at least lam many, and they form a set of branches whose nodes are the wide nodes extending p together with the truncations of p itself. The tree bound therefore caps them by ded of that node set, so if the wide nodes extending p were at most ฮผ the count would be at most ded (ฮผ + #W), which is below lam.
โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ]
{F : Set (Lex (W โ ฮฒ))} {lam ฮผ : Cardinal.{u}} {w : W} {p : Lex (W โ ฮฒ)} [WellFoundedLT W],
lam.IsRegular โ
Cardinal.mk โ(F \ coreโ lam F) < lam โ
p โ bigAtโ lam F w โ ded (ฮผ + Cardinal.mk W) < lam โ ฮผ < Cardinal.mk โ{q | q โ bigโ lam F โง truncate w q = p}โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ] (w v : W)
(b : Lex (W โ ฮฒ)), truncate w (truncate v b) = truncate (min w v) bThe tree bound: a set ๐ฎ of sequences all of whose truncations lie in ๐ has at most ded #๐ elements. Equivalently, a tree with at most ฮบ nodes has at most ded ฮบ branches. The order is taken on ๐ฎ โช ๐ with ๐ as the dense subset: given f < g there, either g lies in ๐ and brackets the pair by itself, or g lies in ๐ฎ and the truncation of g at the first index where f and g differ lies between them.
โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ] [WellFoundedLT W]
{๐ฎ ๐ : Set (Lex (W โ ฮฒ))} {ฮบ : Cardinal.{u}},
(โ f โ ๐ฎ, โ (w : W), truncate w f โ ๐) โ Cardinal.mk โ๐ โค ฮบ โ Cardinal.mk โ๐ฎ โค ded ฮบโ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ] [WellFoundedLT W]
(w : W) (b : Lex (W โ ฮฒ)), truncate w b โค bThe defining property of ded, in the form used downstream: a linear order with a dense subset of size at most ฮบ has at most ded ฮบ points.
โ {ฮบ : Cardinal.{u}} {J : Type u} [inst : LinearOrder J] {E : Set J},
Cardinal.mk โE โค ฮบ โ IsDenseIn E โ Cardinal.mk J โค ded ฮบโ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ] {a b : Lex (W โ ฮฒ)}
{w : W}, (โ j < w, ofLex a j = ofLex b j) โ ofLex a w < ofLex b w โ a < truncate w bโ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ] (w : W)
(b : Lex (W โ ฮฒ)) (v : W), ofLex (truncate w b) v = if v โค w then ofLex b v else โฅโ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] {a b : Lex (W โ ฮฒ)},
a < b โ โ i, (โ j < i, ofLex a j = ofLex b j) โง ofLex a i < ofLex b iA wide node stays wide inside the core, the pruning having removed fewer than lam members.
โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ]
{F : Set (Lex (W โ ฮฒ))} {lam : Cardinal.{u}} {w : W} {p : Lex (W โ ฮฒ)},
lam.IsRegular โ
Cardinal.mk โ(F \ coreโ lam F) < lam โ
p โ bigAtโ lam F w โ lam โค Cardinal.mk โ{g | g โ coreโ lam F โง truncate w g = p}The wide nodes below a fixed node spread over the levels, so if they are more than ฮผ in total and the levels are at most ฮผ, one level already carries more than ฮผ of them.
โ {W : Type u} [inst : LinearOrder W] {ฮฒ : Type u} [inst_1 : LinearOrder ฮฒ] [inst_2 : OrderBot ฮฒ]
{F : Set (Lex (W โ ฮฒ))} {lam ฮผ : Cardinal.{u}} {w : W} {p : Lex (W โ ฮฒ)},
Cardinal.aleph0 โค ฮผ โ
Cardinal.mk W โค ฮผ โ
ฮผ < Cardinal.mk โ{q | q โ bigโ lam F โง truncate w q = p} โ
โ v, ฮผ < Cardinal.mk โ{q | q โ bigAtโ lam F v โง truncate w q = p}โ {ฮฑ W : Type u} {e : W โ ฮฑ} (C : Set ฮฑ) (w : W), ofLex (encโ e C) w = valInโ โจ ofLex (encโ e C) w = valOutโโ {ฮฑ W : Type u} {e : W โ ฮฑ} (C : Set ฮฑ) (w : W), ofLex (encโ e C) w = valInโ โ e w โ CvalInโ โ valOutโ
โ {ฮฑ W : Type u} {e : W โ ฮฑ} {C : Set ฮฑ} (w : W), ofLex (encโ e C) w = if e w โ C then valInโ else valOutโThe reduction that makes the tree small. Among the subsets of A whose trace family outruns ฮธ there is one of least cardinality, Cardinal being well-ordered. Along an enumeration of that subset every proper initial segment carries at most ฮธ traces, which is what the tree bound needs and what a well-order of A itself does not supply.
โ {ฮฑ : Type u} {๐ : Set (Set ฮฑ)} {A : Set ฮฑ} {ฮธ : Cardinal.{u}},
ฮธ < Cardinal.mk โ((fun x => A โฉ x) '' ๐) โ
โ B โ A,
ฮธ < Cardinal.mk โ((fun x => B โฉ x) '' ๐) โง
โ B' โ A, Cardinal.mk โB' < Cardinal.mk โB โ Cardinal.mk โ((fun x => B' โฉ x) '' ๐) โค ฮธโ {ฮบโ ฮบโ : Cardinal.{u}}, ฮบโ โค ฮบโ โ ded ฮบโ โค ded ฮบโโ (ฮบ : Cardinal.{u}), ฮบ โ dedSetโ ฮบโ {I : Type u_1} [inst : Preorder I], IsDenseIn Set.univโ (ฮบ : Cardinal.{u}), BddAbove (dedSetโ ฮบ)โ {ฮบ c : Cardinal.{u}}, c โ dedSetโ ฮบ โ c โค 2 ^ (ฮบ + ฮบ)A linear order with a dense subset D has at most 2 ^ #D * 2 ^ #D points, since a point is determined by the pair of cuts it induces on D.
โ {J : Type u} [inst : LinearOrder J] {E : Set J}, IsDenseIn E โ Cardinal.mk J โค 2 ^ Cardinal.mk โE * 2 ^ Cardinal.mk โEโ {I : Type u_1} {D : Set I} [inst : LinearOrder I], IsDenseIn D โ Function.Injective (cutCodeโ D)โ {I : Type u_1} {D : Set I} [inst : LinearOrder I], IsDenseIn D โ โ {x y : I}, x < y โ cutCodeโ D x โ cutCodeโ D yHasVCDimLE d ๐
A.Infinite
A set family ๐ has VC dimension at most d if all the sets it shatters have size at most d.
{ฮฑ : Type u_1} โ โ โ Set (Set ฮฑ) โ PropD is dense in a preorder I when every pair a < b brackets a point of D, that is, when there is d โ D with a โค d โค b. On a linear order without jumps this agrees with strict betweenness; on an order with jumps it is strictly weaker, and it is the form under which ded bounds the size of the order.
{I : Type u_1} โ [Preorder I] โ Set I โ PropA set family ๐ shatters a set A if all subsets of A can be obtained as the intersection of A with some element of the set family. We also say that A is traced by ๐.
{ฮฑ : Type u_1} โ [SemilatticeInf ฮฑ] โ Set ฮฑ โ ฮฑ โ PropThe wide nodes of the tree of truncations of F, at all levels.
{W : Type u} โ
[LinearOrder W] โ
{ฮฒ : Type u} โ [inst : LinearOrder ฮฒ] โ [OrderBot ฮฒ] โ Cardinal.{u} โ Set (Lex (W โ ฮฒ)) โ Set (Lex (W โ ฮฒ))The level-w truncations of F that at least lam members of F extend.
{W : Type u} โ
[LinearOrder W] โ
{ฮฒ : Type u} โ [inst : LinearOrder ฮฒ] โ [OrderBot ฮฒ] โ Cardinal.{u} โ Set (Lex (W โ ฮฒ)) โ W โ Set (Lex (W โ ฮฒ))The members of F all of whose truncations are wide.
{W : Type u} โ
[LinearOrder W] โ
{ฮฒ : Type u} โ [inst : LinearOrder ฮฒ] โ [OrderBot ฮฒ] โ Cardinal.{u} โ Set (Lex (W โ ฮฒ)) โ Set (Lex (W โ ฮฒ))The pair of cuts a point of I induces on D.
{I : Type u_1} โ [Preorder I] โ (D : Set I) โ I โ Set โD ร Set โDded ฮบ is the supremum of the cardinalities of the linear orders that admit a dense subset of size at most ฮบ. The supremum is taken over a set of cardinals bounded above by 2 ^ (ฮบ + ฮบ), so it is not a junk value. Carriers are restricted to Type u, which does not move the supremum, every witness having size at most 2 ^ (ฮบ + ฮบ).
Cardinal.{u} โ Cardinal.{u}The set of cardinalities of linear orders with a dense subset of size at most ฮบ.
Cardinal.{u} โ Set Cardinal.{u}The code of C along an indexing e of the ground set.
{ฮฑ W : Type u} โ (W โ ฮฑ) โ Set ฮฑ โ Lex (W โ WithBot (ULift.{u, 0} Bool))The sequence agreeing with b up to w and equal to โฅ above it.
{W : Type u} โ [LinearOrder W] โ {ฮฒ : Type u} โ [inst : LinearOrder ฮฒ] โ [OrderBot ฮฒ] โ W โ Lex (W โ ฮฒ) โ Lex (W โ ฮฒ)The two defined values of the alphabet coding a partial trace; โฅ marks the positions where the trace is undefined.
WithBot (ULift.{u, 0} Bool)WithBot (ULift.{u, 0} Bool)No cardinal function strictly smaller than ded bounds the traces of a family of finite VC dimension. For every infinite ฮบ and every lam < ded ฮบ there is a family of VC dimension 1 on a ground set of size exactly ฮบ tracing more than lam sets. Together with HasVCDimLE.mk_image_inter_le_ded, which caps those traces at ded #A, this places ded at the exact strength of the bound.
The conclusion cannot be strengthened to a family tracing exactly ded ฮบ sets. Chernikov and Shelah, On the number of Dedekind cuts and two-cardinal models of dependent theories, note after their definition of ded ฮบ that "in general the supremum need not be attained", so at such a ฮบ no family attains it and quantifying below the supremum is the available form. At ฮบ = โตโ the supremum is attained, by mk_image_inter_range_cut_eq_ded.
โ {ฮบ : Cardinal.{u}},
Cardinal.aleph0 โค ฮบ โ
โ {lam : Cardinal.{u}},
lam < ded ฮบ โ โ ฮฑ ๐ A, Cardinal.mk โA = ฮบ โง HasVCDimLE 1 ๐ โง lam < Cardinal.mk โ((fun x => A โฉ x) '' ๐)No cardinal function smaller than ded bounds the traces, which places ded at the exact strength of the bound.
Asserted byโ {ฮบ : Cardinal.{u}},
Cardinal.aleph0 โค ฮบ โ
โ {lam : Cardinal.{u}},
lam < ded ฮบ โ โ ฮฑ ๐ A, Cardinal.mk โA = ฮบ โง HasVCDimLE 1 ๐ โง lam < Cardinal.mk โ((fun x => A โฉ x) '' ๐)โ (ฮบ : Cardinal.{u}), 1 โ dedSet'โ ฮบIsStrictlyDenseIn โ
The witness attached to a strictly dense subset: a family of VC dimension 1 on a ground set of size exactly ฮบ whose traces are at least as many as the points of J.
โ {ฮบ : Cardinal.{u}} {J : Type u} [inst : LinearOrder J] {E : Set J},
Cardinal.aleph0 โค ฮบ โ
Cardinal.mk โE โค ฮบ โ
IsStrictlyDenseIn E โ
โ ฮฑ ๐ A, Cardinal.mk โA = ฮบ โง HasVCDimLE 1 ๐ โง Cardinal.mk J โค Cardinal.mk โ((fun x => A โฉ x) '' ๐)โ {ฮบ : Cardinal.{u}} {J P : Type u} {E : Set J},
Cardinal.aleph0 โค ฮบ โ Cardinal.mk โE โค ฮบ โ Cardinal.mk P = ฮบ โ Cardinal.mk โ(padGroundโ E) = ฮบโ (J P : Type u) [inst : LinearOrder J], IsChain (fun x1 x2 => x1 โ x2) (padCutsโ J P)
A point of E strictly between j and j' separates the two traces: it lies in the trace at j' and, being above j, not in the trace at j.
โ {J P : Type u} [inst : LinearOrder J] {E : Set J} {j j' e : J},
e โ E โ j < e โ e < j' โ padGroundโ E โฉ Sum.inl '' {a | a < j} โ padGroundโ E โฉ Sum.inl '' {a | a < j'}A family linearly ordered by inclusion has VC dimension at most 1. Shattering a two point set asks for a member meeting it in the first point alone and a member meeting it in the second point alone; whichever of the two members is contained in the other already contains its own point, so it meets the set in both. This is why the extremal families for ded are families of cuts: the number of cuts is exactly what a chain in a linear order can achieve.
โ {ฮฑ : Type u_2} {๐ : Set (Set ฮฑ)}, IsChain (fun x1 x2 => x1 โ x2) ๐ โ HasVCDimLE 1 ๐The convention question is empty at infinite cardinals: ded and ded' agree there. It is not empty in general, by ded'_one_lt_ded_one.
โ {ฮบ : Cardinal.{u}}, Cardinal.aleph0 โค ฮบ โ ded' ฮบ = ded ฮบThe two readings of "dense subset" define the same cardinal function at every infinite cardinal. Every bracketing witness is turned into a strict witness of at least its size by mk_le_ded'_of_isDenseIn, at the cost of a countable factor in the dense subset, which an infinite ฮบ absorbs.
โ {ฮบ : Cardinal.{u}}, Cardinal.aleph0 โค ฮบ โ ded ฮบ โค ded' ฮบโ {ฮบ : Cardinal.{u}} {J : Type u} [inst : LinearOrder J] {E : Set J},
Cardinal.aleph0 โค ฮบ โ Cardinal.mk โE โค ฮบ โ IsDenseIn E โ Cardinal.mk J โค ded' ฮบโ {J : Type u} (E : Set J), Cardinal.mk J โค Cardinal.mk โ(fattenโ E)โ {J : Type u} (E : Set J), Function.Injective fun x => โจtoLex (x, 0), โฏโฉThe defining property of ded': a linear order with a strictly dense subset of size at most ฮบ has at most ded' ฮบ points.
โ {ฮบ : Cardinal.{u}} {J : Type u} [inst : LinearOrder J] {E : Set J},
Cardinal.mk โE โค ฮบ โ IsStrictlyDenseIn E โ Cardinal.mk J โค ded' ฮบโ (ฮบ : Cardinal.{u}), BddAbove (dedSet'โ ฮบ)โ {ฮบ c : Cardinal.{u}}, c โ dedSet'โ ฮบ โ c โค 2 ^ (ฮบ + ฮบ)โ {J : Type u} (E : Set J), Cardinal.mk โ(fattenDenseโ E) โค Cardinal.mk โE * Cardinal.aleph0โ {J : Type u} (E : Set J), Function.Injective fun p => (โจ(ofLex โโp).1, โฏโฉ, (ofLex โโp).2)โ {J : Type u} [inst : LinearOrder J] {E : Set J}, IsDenseIn E โ IsStrictlyDenseIn (fattenDenseโ E)The heart of the comparison: in the blow-up, every pair is separated strictly by a point lying over E. The three cases of the bracketing witness d are a < d < b, where the fibre over d supplies the point; d = a, where the point is taken higher in the fibre over a; and d = b, where it is taken lower in the fibre over b.
โ {J : Type u} [inst : LinearOrder J] {E : Set J},
IsDenseIn E โ
โ {u v : Lex (J ร โ)}, u โ fattenโ E โ v โ fattenโ E โ u < v โ โ z โ E, โ r, u < toLex (z, r) โง toLex (z, r) < vโ {J : Type u} {E : Set J} {u v : Lex (J ร โ)},
u โ fattenโ E โ v โ fattenโ E โ (ofLex u).1 = (ofLex v).1 โ (ofLex u).2 < (ofLex v).2 โ (ofLex u).1 โ EThe elimination rule for ded.
โ {ฮบ c : Cardinal.{u}},
(โ (J : Type u) (x : LinearOrder J) (E : Set J), Cardinal.mk โE โค ฮบ โ IsDenseIn E โ Cardinal.mk J โค c) โ ded ฮบ โค cStrict density is the stronger condition on the subset, so fewer orders are witnesses and the supremum is no larger. This holds at every cardinal.
โ (ฮบ : Cardinal.{u}), ded' ฮบ โค ded ฮบThe defining property of ded, in the form used downstream: a linear order with a dense subset of size at most ฮบ has at most ded ฮบ points.
โ {ฮบ : Cardinal.{u}} {J : Type u} [inst : LinearOrder J] {E : Set J},
Cardinal.mk โE โค ฮบ โ IsDenseIn E โ Cardinal.mk J โค ded ฮบโ (ฮบ : Cardinal.{u}), BddAbove (dedSetโ ฮบ)โ {ฮบ c : Cardinal.{u}}, c โ dedSetโ ฮบ โ c โค 2 ^ (ฮบ + ฮบ)A linear order with a dense subset D has at most 2 ^ #D * 2 ^ #D points, since a point is determined by the pair of cuts it induces on D.
โ {J : Type u} [inst : LinearOrder J] {E : Set J}, IsDenseIn E โ Cardinal.mk J โค 2 ^ Cardinal.mk โE * 2 ^ Cardinal.mk โEโ {I : Type u_1} {D : Set I} [inst : LinearOrder I], IsDenseIn D โ Function.Injective (cutCodeโ D)โ {I : Type u_1} {D : Set I} [inst : LinearOrder I], IsDenseIn D โ โ {x y : I}, x < y โ cutCodeโ D x โ cutCodeโ D yThe elimination rule for ded'.
โ {ฮบ c : Cardinal.{u}},
(โ (J : Type u) (x : LinearOrder J) (E : Set J), Cardinal.mk โE โค ฮบ โ IsStrictlyDenseIn E โ Cardinal.mk J โค c) โ
ded' ฮบ โค cโ {I : Type u_1} {D : Set I} [inst : Preorder I], IsStrictlyDenseIn D โ IsDenseIn DCardinal.aleph0 โค ฮบ
lam < ded ฮบ
A set family ๐ has VC dimension at most d if all the sets it shatters have size at most d.
{ฮฑ : Type u_1} โ โ โ Set (Set ฮฑ) โ PropD is dense in a preorder I when every pair a < b brackets a point of D, that is, when there is d โ D with a โค d โค b. On a linear order without jumps this agrees with strict betweenness; on an order with jumps it is strictly weaker, and it is the form under which ded bounds the size of the order.
{I : Type u_1} โ [Preorder I] โ Set I โ PropD is strictly dense in a preorder I when every pair a < b has a point of D strictly between them. This is the second of the two readings of "dense subset" that the definition of ded in the literature leaves open; the first is IsDenseIn. Where the order has no jumps the two agree, and a linear order with a jump whose two endpoints lie in D satisfies IsDenseIn only. The literature on ded does not say which reading is meant; ded'_eq_ded shows the question does not affect the value of ded at an infinite cardinal.
{I : Type u_1} โ [Preorder I] โ Set I โ PropA set family ๐ shatters a set A if all subsets of A can be obtained as the intersection of A with some element of the set family. We also say that A is traced by ๐.
{ฮฑ : Type u_1} โ [SemilatticeInf ฮฑ] โ Set ฮฑ โ ฮฑ โ PropThe pair of cuts a point of I induces on D.
{I : Type u_1} โ [Preorder I] โ (D : Set I) โ I โ Set โD ร Set โDded ฮบ is the supremum of the cardinalities of the linear orders that admit a dense subset of size at most ฮบ. The supremum is taken over a set of cardinals bounded above by 2 ^ (ฮบ + ฮบ), so it is not a junk value. Carriers are restricted to Type u, which does not move the supremum, every witness having size at most 2 ^ (ฮบ + ฮบ).
Cardinal.{u} โ Cardinal.{u}ded' ฮบ is the supremum of the cardinalities of the linear orders that admit a strictly dense subset of size at most ฮบ. It is the same construction as ded, with strict betweenness in place of bracketing, and is bounded above by the same 2 ^ (ฮบ + ฮบ).
Cardinal.{u} โ Cardinal.{u}The set of cardinalities of linear orders with a dense subset of size at most ฮบ.
Cardinal.{u} โ Set Cardinal.{u}The set of cardinalities of linear orders with a strictly dense subset of size at most ฮบ.
Cardinal.{u} โ Set Cardinal.{u}The blow-up of J along E: the lexicographic product J รโ โ with the fibre over each point outside E collapsed to the single rational 0.
{J : Type u} โ Set J โ Set (Lex (J ร โ))The part of the blow-up lying over E, that is, the union of the full โ fibres.
{J : Type u} โ (E : Set J) โ Set โ(fattenโ E)The witness family: the down sets of J, carried into the padded ground type.
(J P : Type u) โ [LinearOrder J] โ Set (Set (J โ P))
The ground set of the witness: a subset E of J together with a disjoint copy of P, whose points no member of the family below meets.
{J P : Type u} โ Set J โ Set (J โ P)โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k
The forward direction of the Moran-Yehudayoff theorem: finite VC dimension implies existence of a compression scheme with finite side information.
The construction: 1. Build a proper finite-support learner L from VC + Sauer-Shelah 2. For sample S: extract c, Y = pointSupport S, HY = hypothesis envelope 3. Apply approximate minimax on the agreement game โ distribution p on HY 4. Apply VC ฮต-approximation on agreement tests โ T representative hypotheses 5. Kernel = union of witness subsets for T hypotheses 6. Side info = incidence: which hypothesis's witness contains each kernel point 7. Reconstruct by majority vote over T hypotheses
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k
Finite VC dimension implies existence of a proper finite-support learner. The construction uses ERM + finite_support_vc_approx on the disagreement family.
โ (X : Type u) (C : ConceptClass X Bool), Set.Nonempty C โ VCDim X C < โค โ โ _L, True
supportError expressed in terms of boolTestExpectation of a disagreement test.
โ {X : Type u} (Y : Finset X) (q : FinitePMF โฅY) (h c : X โ Bool),
supportError Y q h c = boolTestExpectation q fun y => decide (h โy โ c โy)VC dimension of the disagreement family is bounded by VCDim(C). Restriction to Y and xor with c do not increase shattering dimension.
โ {X : Type u} [inst : DecidableEq X] (C : ConceptClass X Bool) (c : X โ Bool) (Y : Finset X) {d : โ},
VCDim X C โค โd โ (disagreementFamilyโ C c Y).boolVCDim โค dThe Moran-Yehudayoff forward construction. Uses finalizeIncidenceScheme to package the majority-vote scheme with universe-correct Info type.
The agent must provide: compressCore, blockHyp, rowHyp, hsmall, hsub, hagree, hmajor. These are the MY wiring.
โ (X : Type u) (C : ConceptClass X Bool),
Set.Nonempty C โ
โ (L : ProperFiniteSupportLearner X C), VCDim X C < โค โ โ (_K : โ), โ k cs, CompressionSchemeWithInfo.size cs = kGeneric roundtrip theorem for the hround sorry.
If:
encodeWitnessInfo is used in compressCore,decodeWitnessXCoords and decodeWitnessLabel are used in blockHyp, andthen the decoded block hypothesis is exactly the representative hypothesis.
โ {X : Type u} [inst : DecidableEq X] (learn : {m : โ} โ (Fin m โ X ร Bool) โ X โ Bool) (kernel : Finset (X ร Bool))
(c : X โ Bool) (K : โ) (W : Finset X) (h : X โ Bool),
kernel.card โค K โ
(โ x โ W, (x, c x) โ kernel) โ
(โ p โ kernel, p.2 = c p.1) โ
learn (labeledSampleOfFinset c W) = h โ
โ (x : X),
have info := encodeWitnessInfo kernel c K W;
have blockXCoords := decodeWitnessXCoords kernel info;
have blockLabel := decodeWitnessLabel kernel;
learn (labeledSampleOfFinset blockLabel blockXCoords) x = h xIf two label functions agree on all points of Z, then the labeled samples they induce on Z.equivFin are equal.
โ {X : Type u} [DecidableEq X] {โโ โโ : X โ Bool} {Z : Finset X},
(โ x โ Z, โโ x = โโ x) โ labeledSampleOfFinset โโ Z = labeledSampleOfFinset โโ ZIf every (x, c x) with x โ W lies in kernel, and kernel.card โค K, then decoding the encoded witness positions gives back exactly W.
โ {X : Type u} [inst : DecidableEq X] (kernel : Finset (X ร Bool)) (c : X โ Bool) {K : โ} (W : Finset X),
kernel.card โค K โ (โ x โ W, (x, c x) โ kernel) โ decodeWitnessXCoords kernel (encodeWitnessInfo kernel c K W) = WOn the encoded witness support, the decoded label function agrees with the true label function c, provided every pair in the kernel has the correct second coordinate.
โ {X : Type u} [inst : DecidableEq X] (kernel : Finset (X ร Bool)) (c : X โ Bool) (W : Finset X),
(โ x โ W, (x, c x) โ kernel) โ (โ p โ kernel, p.2 = c p.1) โ โ x โ W, decodeWitnessLabel kernel x = c xGenuine approximate minimax via MWU regret extraction. If every column mixture admits a pure row with expected payoff โฅ v, then there is a row mixture with payoff โฅ v - ฮต against every column.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [Nonempty R] [Nonempty C] [DecidableEq R]
[DecidableEq C] (M : R โ C โ Bool) (v ฮต : โ),
0 < ฮต โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ
โ p, โ (c : C), v - ฮต โค boolGamePayoff M p cA single weight is bounded by the potential.
โ {C : Type u_1} [inst : Fintype C] (cfg : MWUConfig C) (c : C), cfg.weights c โค cfg.potentialExact individual-weight tracking: the weight of column c after T rounds is (1-ฮท) to the number of rounds in which c was hit.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] (M : R โ C โ Bool) (ฮท : โ)
(hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) (T : โ)
(c : C), (mwuConfig M ฮท hฮท1 v hrow T).weights c = (1 - ฮท) ^ mwuHitCountโ M ฮท hฮท1 v hrow T cโ {C : Type u_1} [inst : Fintype C] (weights weights_1 : C โ โ) (e_weights : weights = weights_1)
(weights_pos : โ (c : C), 0 < weights c),
{ weights := weights, weights_pos := weights_pos } = { weights := weights_1, weights_pos := โฏ }Potential bound after T steps: ฮฆ_T โค |C| ยท (1 - ฮทv)^T.
This is the core MWU guarantee. Combined with individual weight lower bounds (w_T(c) = (1-ฮท)^{losses(c)}), it yields the regret bound.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] (M : R โ C โ Bool)
(ฮท : โ),
0 โค ฮท โ
โ (hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0)
(T : โ), (mwuConfig M ฮท hฮท1 v hrow T).potential โค โ(Fintype.card C) * (1 - ฮท * v) ^ TPotential bound after one step: ฮฆ' โค ฮฆ ยท (1 - ฮทยทv).
โ {R : Type u_1} {C : Type u_2} [Fintype R] [inst : Fintype C] [inst_1 : Nonempty C] (M : R โ C โ Bool) (ฮท : โ),
0 โค ฮท โ
โ (hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0)
(cfg : MWUConfig C), (mwuUpdateWeights M ฮท hฮท1 cfg โฏ.choose).potential โค cfg.potential * (1 - ฮท * v)Best response payoff โฅ v ยท ฮฆ in terms of weights.
โ {R : Type u_1} {C : Type u_2} [Fintype R] [inst : Fintype C] [inst_1 : Nonempty C] (M : R โ C โ Bool) (v : โ)
(hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) (cfg : MWUConfig C),
v * cfg.potential โค โ c, cfg.weights c * if M โฏ.choose c = true then 1 else 0โ {C : Type u_1} [inst : Fintype C] [Nonempty C] (cfg : MWUConfig C), 0 < cfg.potentialโ {C : Type u_1} [inst : Fintype C] (self : MWUConfig C) (c : C), 0 < self.weights cโ {C : Type u_1} [inst : Fintype C] {R : Type u_2} (M M_1 : R โ C โ Bool),
M = M_1 โ
โ (ฮท ฮท_1 : โ) (e_ฮท : ฮท = ฮท_1) (hฮท1 : ฮท < 1) (cfg cfg_1 : MWUConfig C),
cfg = cfg_1 โ โ (r r_1 : R), r = r_1 โ mwuUpdateWeights M ฮท hฮท1 cfg r = mwuUpdateWeights M_1 ฮท_1 โฏ cfg_1 r_1โ (C : Type u_1) [inst : Fintype C], (mwuInit C).potential = โ(Fintype.card C)
The minimax value of a Boolean game is at most 1.
โ {R : Type u_1} {C : Type u_2} [Fintype R] [inst : Fintype C] [Nonempty C] (M : R โ C โ Bool) (v : โ),
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ v โค 1Arithmetic core: from the potential bound and sufficiently small ฮท / large T, deduce a per-column hit-rate lower bound. Uses Real.log โ exactly 4 Mathlib lemmas.
โ {N H T : โ} {ฮท v ฮต : โ},
0 < โN โ
0 < ฮท โ
ฮท < 1 โ
v โค 1 โ
0 < T โ (1 - ฮท) ^ H โค โN * (1 - ฮท * v) ^ T โ ฮท โค ฮต / 4 โ Real.log โN / (ฮท * โT) โค ฮต / 4 โ v - ฮต โค โH / โTโ {R : Type u_1} {C : Type u_2} [inst : Fintype R] (M : R โ C โ Bool) (p : FinitePMF R) (c : C),
0 โค boolGamePayoff M p cEmpirical payoff of the MWU row sequence equals the normalized hit count.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] [inst_3 : DecidableEq R]
(M : R โ C โ Bool) (ฮท : โ) (hฮท1 : ฮท < 1) (v : โ)
(hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) {T : โ} (hT : 0 < T) (c : C),
boolGamePayoff M (empiricalPMF hT (mwuRows M ฮท hฮท1 v hrow T)) c = โ(mwuHitCountโ M ฮท hฮท1 v hrow T c) / โTThe recursive hit counter agrees with the sum of Boolean indicators over the emitted row sequence.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] (M : R โ C โ Bool) (ฮท : โ)
(hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) (T : โ)
(c : C), โ(mwuHitCountโ M ฮท hฮท1 v hrow T c) = โ t, if M (mwuRows M ฮท hฮท1 v hrow T t) c = true then 1 else 0Specialized empirical-payoff identity for ApproxMinimax (avoids cyclic import with FiniteVCApprox).
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : DecidableEq R] {T : โ} (hT : 0 < T) (rs : Fin T โ R)
(M : R โ C โ Bool) (c : C), boolGamePayoff M (empiricalPMF hT rs) c = (โ t, if M (rs t) c = true then 1 else 0) / โTEvery hypothesis in the envelope is in C.
โ {X : Type u} {C : ConceptClass X Bool} (L : ProperFiniteSupportLearner X C) (c : X โ Bool) (Y : Finset X),
โ h โ hypothesisEnvelope L c Y, h โ Cโ {X : Type u} {C : ConceptClass X Bool} (self : ProperFiniteSupportLearner X C) {m : โ} (S : Fin m โ X ร Bool),
self.learn S โ CFor each C-realizable sample, the proper learner provides a row-response for the minimax game on the hypothesis envelope.
โ {X : Type u} {C : ConceptClass X Bool} (L : ProperFiniteSupportLearner X C),
โ c โ C,
โ (Y : Finset X) [Nonempty โฅY] (HY : Finset (X โ Bool)),
HY = hypothesisEnvelope L c Y โ
โ (q : FinitePMF โฅY), โ h, 2 / 3 โค โ y, q.prob y * if decide (โh โy = c โy) = true then 1 else 0Weighted agreement = 1 - supportError.
โ {X : Type u} (Y : Finset X) (q : FinitePMF โฅY) (h c : X โ Bool),
(โ y, q.prob y * if h โy = c โy then 1 else 0) = 1 - supportError Y q h cโ {X : Type u} {C : ConceptClass X Bool} (self : ProperFiniteSupportLearner X C),
โ c โ C,
โ (Y : Finset X) (q : FinitePMF โฅY),
โ Z โ Y, Z.card โค self.sampleBound โง supportError Y q (self.learn (labeledSampleOfFinset c Z)) c โค 1 / 3Finite-support distributions uniformly approximate any distribution on a VC class. For a class of VC dimension at most d and any ฮต > 0, there exists T = T(d, ฮต) such that every finitely supported distribution ฮผ is within ฮต (uniformly over the class) of some empirical distribution on T points. A density-style reduction that lets the approximate minimax / MWU machinery, which lives in finite support, apply to general distributions.
โ (d : โ) (ฮต : โ),
0 < ฮต โ
โ T,
โ (hT : 0 < T),
โ {H : Type u_1} [inst : Fintype H] [inst_1 : DecidableEq H] (A : Finset (H โ Bool)),
A.boolVCDim โค d โ
โ (ฮผ : FinitePMF H),
โ hs, โ a โ A, |boolTestExpectation ฮผ a - boolTestExpectation (empiricalPMF hT hs) a| โค ฮตโ {H : Type u_1} [inst : Fintype H] [DecidableEq H] [inst_2 : MeasurableSpace H] [MeasurableSingletonClass H]
(ฮผ : FinitePMF H) (a : H โ Bool),
TrueErrorReal (H โ โ) (extendBoolโ a) (fun x => false) (PMF.map Sum.inl (FinitePMF.toPMFโ ฮผ)).toMeasure =
boolTestExpectation ฮผ aThe symmetrization uniform convergence bound: two-sided version. P[โhโC: |TrueErr-EmpErr| โฅ ฮต] โค 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8).
Proof strategy (4 steps):
1. Decompose absolute value: |TrueErr - EmpErr| โฅ ฮต โ (TrueErr - EmpErr โฅ ฮต) โจ (EmpErr - TrueErr โฅ ฮต)
have abs_decomp : โ (a b : โ),
|a - b| โฅ ฮต โ a - b โฅ ฮต โจ b - a โฅ ฮต := by
intro a b; constructor
ยท intro h; by_cases h' : a - b โฅ ฮต
ยท exact Or.inl h'
ยท exact Or.inr (by linarith [abs_sub_comm a b, le_abs_self (a - b)])
ยท intro h; cases h with
| inl h => exact le_trans (le_of_eq (abs_of_nonneg (by linarith))) (by linarith)
| inr h => exact le_trans (le_of_eq (abs_of_nonpos (by linarith) โธ ...)) ...2. Upper tail: P[โhโC: TrueErr-EmpErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step + double_sample_pattern_bound.3. Lower tail: P[โhโC: EmpErr-TrueErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step to the event EmpErr-TrueErr โฅ ฮต and bound the double-sample event {EmpErr_S - EmpErr_{S'} โฅ ฮต/2}. have swap_symmetry : DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S) - EmpErr(S') โฅ ฮต/2} = DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S') - EmpErr(S) โฅ ฮต/2} := Measure.prod_swap ... 4. Union bound: P[|gap| โฅ ฮต] โค P[gap โฅ ฮต] + P[gap โค -ฮต] โค 2ยทGFยทexp(...) + 2ยทGFยทexp(...) = 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
-- Uses: MeasureTheory.measure_union_le for the union of two events
-- CAST: 2 * X + 2 * X = 4 * X in ENNReal (need ENNReal.add_mul or similar)References: SSBD Theorem 6.7, Kakade-Tewari Lecture 19
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
|TrueErrorReal X h c D -
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮต} โค
ENNReal.ofReal (4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))Symmetrization step for the lower tail: P[โh: EmpErr-TrueErr โฅ ฮต] โค 2ยทP_{double}[โh: EmpErr_S-EmpErr_{S'} โฅ ฮต/2].
Mirror of symmetrization_step for the opposite direction. Uses hoeffding_one_sided_upper instead of hoeffding_one_sided.
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) - TrueErrorReal X h c D โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}Upper-tail Hoeffding: for iid Bernoulli(p) draws, the empirical average overshoots the mean by โฅ t with probability โค exp(-2mtยฒ).
This is the mirror of hoeffding_one_sided (which bounds the lower tail). The proof uses the same sub-Gaussian machinery with Z_i = indicator(x_i) - p (instead of p - indicator(x_i)).
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ TrueErrorReal X h c D + t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))Symmetrization: the probability of a large gap TrueErr-EmpErr is at most twice the probability of a large gap EmpErr'-EmpErr on the double sample.
Proof strategy (6 steps):
1. Witness selection: For S in the bad event, โh* โ C with TrueErr(h) - EmpErr_S(h) โฅ ฮต.
-- In the bad event set, extract h* by classical choice
have h_witness : โ xs โ bad_event, โ h* โ C,
TrueErrorReal X h* c D - EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โฅ ฮต2. Ghost sample mean: E_{S'}[EmpErr_{S'}(h)] = TrueErr(h) โฅ EmpErr_S(h*) + ฮต.
MeasureTheory.integral_pi to compute E[EmpErr] over product measure. have expected_emp_err : โ h* : Concept X Bool, โซ xs, EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โ(Measure.pi (fun _ : Fin m => D)) = TrueErrorReal X h* c D := by ... 3. Hoeffding on ghost sample: P_{S'}[EmpErr_{S'}(h) < TrueErr(h) - ฮต/2] โค exp(-mฮตยฒ/2).
hoeffding_one_sided with t = ฮต/2.hm_large hypothesis ensures exp(-mฮตยฒ/2) < 1/2: 2ยทln2 โค mฮตยฒ โน mฮตยฒ/2 โฅ ln2 โน exp(-mฮตยฒ/2) โค 1/2. have hoeffding_ghost : โ h* โ C, Measure.pi (fun _ : Fin m => D) {xs' | EmpiricalError X Bool h* (sample xs') (zeroOneLoss Bool) < TrueErrorReal X h* c D - ฮต/2} โค ENNReal.ofReal (Real.exp (-m * (ฮต/2)^2 * 2)) := by intro h* _; exact hoeffding_one_sided D h* c m hm (ฮต/2) (by linarith) (by ...) (by ...) 4. Complementary probability: P_{S'}[EmpErr_{S'}(h) - EmpErr_S(h) โฅ ฮต/2] โฅ 1/2.
5. Conditional to unconditional: The witness h* from step 1 also witnesses the double-sample event โhโC: EmpErr'-EmpErr โฅ ฮต/2. So: P_{S'}[double event | S bad] โฅ 1/2.
have conditional_bound : โ xs โ bad_event,
Measure.pi (fun _ : Fin m => D)
{xs' | โ h โ C, EmpiricalError ... xs' - EmpiricalError ... xs โฅ ฮต/2}
โฅ ENNReal.ofReal (1/2) := by ...6. Fubini integration: By Measure.prod_apply and Fubini: P_{S,S'}[double event] = โซ_S P_{S'}[double event | S] โฅ (1/2) ยท P_S[bad event] โน P_S[bad event] โค 2 ยท P_{S,S'}[double event].
-- Uses: MeasureTheory.Measure.prod_apply or lintegral_prod
-- MEASURABILITY: the double-sample event is measurable as a finite union
-- of sets of the form {(xs,xs') | EmpErr'(h) - EmpErr(h) โฅ ฮต/2} for h โ C.
-- Since C may be infinite, measurability requires care: the sup over h
-- must be shown to be measurable. For finite restriction patterns (โค 2^m
-- on Fin m โ Bool), this is a finite union.MEASURABILITY CONCERNS:
{xs | โ h โ C, ...} is NOT obviously measurable for infinite C. Strategy: decompose via restriction patterns. On any fixed xs, the set of labelings {(h(xs 0), ..., h(xs(m-1))) | h โ C} has at most GF(C,m) โค 2^m elements. So the โh event is a finite union of measurable sets.EmpiricalError is a finite sum of measurable functions, hence measurable.References: SSBD Lemma 4.5, Kakade-Tewari Lecture 19 Lemma 1
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
TrueErrorReal X h c D - EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}One-sided Hoeffding: for iid Bernoulli(p) draws, the empirical average undershoots the mean by โฅ t with probability โค exp(-2mtยฒ).
Proof strategy (3 steps):
1. MGF bound (Hoeffding's lemma): For X โ [0,1] with E[X] = p, E[exp(s(X-p))] โค exp(sยฒ/8).
cosh_le_exp_sq_half infrastructure in Rademacher.lean. have mgf_bound : โ (s : โ), โซ x, Real.exp (s * (indicator x - p)) โD โค Real.exp (s^2 / 8) := by ... 2. Product independence: E[exp(sยทโ(X_i-p))] = โ E[exp(s(X_i-p))] โค exp(msยฒ/8).
MeasureTheory.Measure.pi independence structure.Measure.pi integral factorization for product of functions.fun xs => Real.exp (s * โ i, f (xs i)) is measurable (composition of measurable functions). have product_bound : โ (s : โ), โซ xs, Real.exp (s * โ i, (indicator (xs i) - p)) โMeasure.pi (fun _ => D) โค Real.exp (m * s^2 / 8) := by ... 3. Exponential Markov + optimize: P[โ(X_i-p) โค -mt] = P[exp(-sยทโ(X_i-p)) โฅ exp(smt)] โค exp(-smt + msยฒ/8). Optimize over s: set s = 4t to get โค exp(-2mtยฒ).
have markov_step : โ (s : โ) (hs : 0 < s), Measure.pi (fun _ => D) {xs | โ i, (indicator (xs i) - p) โค -(m : โ) * t} โค ENNReal.ofReal (Real.exp (-(s * m * t) + m * s^2 / 8)) := by ... have optimize : Real.exp (-(4*t * m * t) + m * (4*t)^2 / 8) = Real.exp (-2 * m * t^2) := by ring_nf CAST ISSUES to watch:
m : โ needs cast to โ in the exponent: (m : โ)EmpiricalError returns โ, TrueErrorReal returns โ, good โ no ENNReal gapENNReal, the bound exp(-2mtยฒ) is โโฅ0โ via ENNReal.ofRealReferences: SSBD Lemma B.3, Hoeffding (1963)
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โค TrueErrorReal X h c D - t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))On the double sample, the probability that any hypothesis has EmpErr' - EmpErr โฅ ฮต/2 is bounded by GF(C,2m) ยท exp(-mฮตยฒ/8).
Proof strategy (Approach A โ standard exchangeability, 5 steps):
1. EXCHANGEABILITY: Under D^m โ D^m, the 2m draws zโ,...,z_{2m} are iid from D. The joint distribution is invariant under permutations of {1,...,2m}.
Key lemma: P_{D^mโD^m}[event(S,S')] = E_z[P_{split}[event | z]] where z = merged sample and the split is uniformly random among all C(2m,m) ways to partition z into two groups of m.
-- Measure.pi permutation invariance
have pi_perm_invariant : โ (ฯ : Equiv.Perm (Fin (2*m))),
(Measure.pi (fun _ : Fin (2*m) => D)).map (fun z i => z (ฯ i))
= Measure.pi (fun _ : Fin (2*m) => D) := by ...
-- Consequence: the event probability equals the split-averaged probability
have exchangeability :
DoubleSampleMeasure D m {p | โ h โ C, gap(p) โฅ ฮต/2}
= โซ z, SplitMeasure m {vs | โ h โ C, gap(split z vs) โฅ ฮต/2}
โ(Measure.pi (fun _ : Fin (2*m) => D)) := by ...2. CONDITIONING: For fixed merged sample z of 2m points:
-- Number of distinct patterns
have num_patterns : โ (z : MergedSample X m),
Set.ncard {p : Fin (2*m) โ Bool | โ h โ C, โ i, p i = (h (z i) โ c (z i))}
โค GrowthFunction X C (2*m) := by ...3. PER-PATTERN HOEFFDING ON SPLITS: For fixed z and fixed pattern p: Under uniformly random split (S,S') of z into two groups of m: diff(p, split) = (1/m) โ_{iโS'} a_i - (1/m) โ_{iโS} a_i
This is a function of the random partition. By Hoeffding's inequality for sampling without replacement (Serfling 1974): P_split[diff โฅ ฮต/2] โค exp(-mฮตยฒ/8)
Alternative derivation: Hoeffding without replacement from Hoeffding with replacement (iid signs) via coupling. The without-replacement bound is actually TIGHTER (variance reduction), but the with-replacement bound suffices.
-- Per-pattern concentration
have per_pattern_bound : โ (z : MergedSample X m) (a : Fin (2*m) โ โ)
(ha : โ i, a i โ Set.Icc 0 1),
SplitMeasure m {vs | (1/m) * โ i โ second_group vs, a i
- (1/m) * โ i โ first_group vs, a i โฅ ฮต/2}
โค ENNReal.ofReal (Real.exp (-(m : โ) * (ฮต/2)^2 / 2)) := by ...
-- Note: m*(ฮต/2)^2/2 = mฮตยฒ/84. UNION BOUND: P_split[โ pattern: diff โฅ ฮต/2 | z] โค (number of patterns) ยท max_pattern P_split[diff โฅ ฮต/2] โค GF(C,2m) ยท exp(-mฮตยฒ/8)
have union_bound : โ (z : MergedSample X m),
SplitMeasure m {vs | โ h โ C, gap(split z vs, h) โฅ ฮต/2}
โค ENNReal.ofReal (GrowthFunction X C (2*m) * Real.exp (-(m : โ) * ฮต^2 / 8))
:= by ...5. INTEGRATE: P_{D^mโD^m}[event] = E_z[P_split[event|z]] (by step 1) โค E_z[GF(C,2m) ยท exp(-mฮตยฒ/8)] (by step 4, pointwise) = GF(C,2m) ยท exp(-mฮตยฒ/8) (bound is independent of z)
-- The bound is a constant, so integrating gives the same constant
-- (using IsProbabilityMeasure for the 2m-fold product)Infrastructure needed:
Fin.sumFinEquiv : Fin m โ Fin n โ Fin (m + n) (available in Mathlib)mergeSamples / splitMergedSample (defined above)SplitMeasure and ValidSplit (defined above)Measure.pi permutation invariance (to be proved or imported)GrowthFunction on 2m points + sauer_shelah_exp_bound from Rademacher.leanMEASURABILITY CONCERNS:
GrowthFunction X C (2*m) is a natural number (deterministic), no measurability issue.References: SSBD Theorem 6.7, Hoeffding (1963), Serfling (1974)
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
ฮต โค 2 โ
Set.Nonempty C โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
have ฮผ := MeasureTheory.Measure.pi fun x => D;
(ฮผ.prod ฮผ)
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ (Y : Type v) {inst : DecidableEq Y} [inst_1 : DecidableEq Y] (a a_1 : Y),
a = a_1 โ โ (a_2 a_3 : Y), a_2 = a_3 โ zeroOneLoss Y a a_2 = zeroOneLoss Y a_1 a_3The number of distinct restriction patterns of C on any n points is at most GF(C,n). For z : Fin n โ X, define patterns(z) = {p : Fin n โ Bool | โ h โ C, โ i, p i = (h(z i) โ c(z i))}. Then patterns(z).ncard โค GrowthFunction X C n by definition of GrowthFunction.
โ {X : Type u} [MeasurableSpace X] [Infinite X] (C : ConceptClass X Bool) (c : Concept X Bool) (n : โ) (z : Fin n โ X),
{p | โ h โ C, โ (i : Fin n), p i = decide (h (z i) โ c (z i))}.ncard โค GrowthFunction X C nRademacher MGF bound.
โ {m : โ},
0 < m โ
โ (a : Fin m โ โ) (c : โ),
0 โค c โ
(โ (i : Fin m), |a i| โค c) โ
โ (t : โ),
0 โค t โ
1 / โ(Fintype.card (SignVector m)) * โ ฯ, Real.exp (t * (1 / โm * โ i, a i * boolToSign (ฯ i))) โค
Real.exp (t ^ 2 * c ^ 2 / (2 * โm))cosh(x) โค exp(xยฒ/2). Standard sub-Gaussian bound.
โ (x : โ), Real.cosh x โค Real.exp (x ^ 2 / 2)
Generic finite exchangeability bound. Given a measure-preserving family of transformations on a probability space, a NullMeasurableSet S, and a pointwise bound on the sum of preimage indicators, conclude ฮฝ(S) โค B.
โ {ฮฉ : Type u_1} {G : Type u_2} [inst : MeasurableSpace ฮฉ] [inst_1 : Fintype G] [Nonempty G]
{ฮฝ : MeasureTheory.Measure ฮฉ} [MeasureTheory.IsProbabilityMeasure ฮฝ] (T : G โ ฮฉ โ ฮฉ) (S : Set ฮฉ),
(โ (g : G), MeasureTheory.MeasurePreserving (T g) ฮฝ ฮฝ) โ
MeasureTheory.NullMeasurableSet S ฮฝ โ
โ (B : ENNReal), (โ (z : ฮฉ), โ g, (T g โปยน' S).indicator 1 z โค B * โ(Fintype.card G)) โ ฮฝ S โค Bโ {X : Type u} [MeasurableSpace X] (C : ConceptClass X Bool) (v : โ),
0 < v โ
โ (m : โ),
0 < m โ
โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
ฮด < 1 โ
(โ (n : โ), v โค n โ GrowthFunction X C n โค โ i โ Finset.range (v + 1), n.choose i) โ
(16 * Real.exp 1 * (โv + 1) / ฮต ^ 2) ^ (v + 1) / ฮด โค โm โ
4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)) โค ฮด โง 2 * Real.log 2 โค โm * ฮต ^ 2Pure combinatorial inequality: โ_{i=0}^d C(m,i) โค (em/d)^d for d โค m, d โฅ 1.
โ (d m : โ), 0 < d โ d โค m โ โ i โ Finset.range (d + 1), โ(m.choose i) โค (Real.exp 1 * โm / โd) ^ d
Key arithmetic lemma for PAC bound: for t > 0, t^d * exp(-t) โค (d+1)!/t. Follows from exp(t) โฅ t^(d+1)/(d+1)! (partial sum of Taylor series).
โ {d : โ} {t : โ}, 0 < t โ t ^ d * Real.exp (-t) โค โ(d + 1).factorial / tTrivial bound: GrowthFunction โค 2^n for all concept classes. Each restriction to an n-element set yields a function in S โ Bool, and there are at most 2^n such functions.
โ {X : Type u} (C : ConceptClass X Bool) (n : โ), GrowthFunction X C n โค 2 ^ nA convex combination of values in {0, 1} is nonnegative.
โ {H : Type u_1} [inst : Fintype H] (ฮผ : FinitePMF H) (f : H โ Bool), 0 โค boolTestExpectation ฮผ fA convex combination of values in {0, 1} is at most 1.
โ {H : Type u_1} [inst : Fintype H] (ฮผ : FinitePMF H) (f : H โ Bool), boolTestExpectation ฮผ f โค 1โ {H : Type u_1} [inst : Fintype H] (self : FinitePMF H) (h : H), 0 โค self.prob hโ {H : Type u_1} [inst : Fintype H] (self : FinitePMF H), โ h, self.prob h = 1Final existential wrapper: closes the theorem in the exact form expected.
โ {X : Type u} {C : ConceptClass X Bool} (T K : โ)
(compressCore : {m : โ} โ (Fin m โ X ร Bool) โ Finset (X ร Bool) ร IncidenceInfo T K)
(blockHyp : Finset (X ร Bool) โ IncidenceInfo T K โ Fin T โ X โ Bool)
(rowHyp : {m : โ} โ (S : Fin m โ X ร Bool) โ (โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) โ Fin T โ X โ Bool),
0 < T โ
(โ {m : โ} (S : Fin m โ X ร Bool), (compressCore S).1.card โค K) โ
(โ {m : โ} (S : Fin m โ X ร Bool), โ(compressCore S).1 โ Set.range S) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m) (t : Fin T),
blockHyp (compressCore S).1 (compressCore S).2 t (S i).1 = rowHyp S hreal t (S i).1) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m),
(โ t, if rowHyp S hreal t (S i).1 = (S i).2 then 1 else 0) / โT > 1 / 2) โ
โ k cs, CompressionSchemeWithInfo.size cs = kโ {X : Type u} {inst : DecidableEq X} [inst_1 : DecidableEq X] (kernel kernel_1 : Finset (X ร Bool)),
kernel = kernel_1 โ
โ (c c_1 : X โ Bool),
c = c_1 โ
โ (K : โ) (W W_1 : Finset X), W = W_1 โ encodeWitnessInfo kernel c K W = encodeWitnessInfo kernel_1 c_1 K W_1Bridges the FinitePMF view and the sample-average view: the expectation of a Bool-valued test under the empirical PMF of a sample equals the sample average (1/T) โ_t f (s_t). This lets the MWU updates and the approximation transfer principle live in the same distributional framework.
โ {H : Type u_1} [inst : Fintype H] [inst_1 : DecidableEq H] {T : โ} (hT : 0 < T) (hs : Fin T โ H) (f : H โ Bool),
boolTestExpectation (empiricalPMF hT hs) f = (โ t, if f (hs t) = true then 1 else 0) / โTIdentifies the game-theoretic payoff (a row distribution against a fixed column in the Bool game) with the corresponding test expectation. The translation that lets the MWU regret bound be applied directly to the compression problem.
โ {R : Type u_1} [inst : Fintype R] [DecidableEq R] {C : Type u_2} (M : R โ C โ Bool) (p : FinitePMF R) (c : C),
boolGamePayoff M p c = boolTestExpectation p fun r => M r cVC dimension of the agreement-test family is bounded by 2^(d+1) - 1, where d bounds the VC dimension of the concept class C. Uses Assouad's coding argument directly: if a shattered set T in โฅHY has |T| โฅ 2^(d+1), embed bitstrings into T, extract d+1 distinct points from Y via shattering, and show these points are shattered by C (using the XOR trick where b(j) = decide(g(x_j) = c(x_j)) absorbs the agree/disagree flip).
โ {X : Type u} [DecidableEq X] (C : ConceptClass X Bool) (c : X โ Bool) (Y : Finset X) (HY : Finset (X โ Bool)),
(โ h โ HY, h โ C) โ โ {d : โ}, VCDim X C โค โd โ (agreeTests c Y HY).boolVCDim โค 2 ^ (d + 1) - 1Compression with side info implies finite VC dimension. Proof by pigeonhole: compress is injective on C-realizable labelings (by correctness), but compressed outputs form a bounded set.
โ (X : Type u) (C : ConceptClass X Bool), (โ k cs, cs.size = k) โ VCDim X C < โค
โ {X : Type u} {C : ConceptClass X Bool} {S T : Finset X}, T โ S โ Shatters X C S โ Shatters X C TExponential beats polynomial for the compression pigeonhole argument.
โ (s : โ), (s + 1) ^ 2 * (4 * (s + 1) ^ 2) ^ s < 2 ^ (2 * (s + 1) * (s + 1))
โ (k : โ), k + 1 โค 2 ^ k
Pigeonhole core: if two C-realizable samples over the same points with different labelings produce the same (kernel, info) pair, correctness forces the labelings to agree.
โ {X : Type u} {n : โ} {C : ConceptClass X Bool} (cs : CompressionSchemeWithInfo X Bool C) (pts : Fin n โ X),
Function.Injective pts โ
โ (f g : Fin n โ Bool),
(โ c โ C, โ (i : Fin n), c (pts i) = f i) โ
(โ c โ C, โ (i : Fin n), c (pts i) = g i) โ
((cs.compress fun i => (pts i, f i)) = cs.compress fun i => (pts i, g i)) โ f = gCorrectness: reconstructed hypothesis agrees with every sample point, when the sample is C-realizable
โ {X : Type u} {Y : Type v} {C : ConceptClass X Y} (self : CompressionSchemeWithInfo X Y C) {m : โ} (S : Fin m โ X ร Y),
(โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) โ
โ (i : Fin m), self.reconstruct (self.compress S).1 (self.compress S).2 (S i).1 = (S i).2Compressed set is a subset of the sample
โ {X : Type u} {Y : Type v} {C : ConceptClass X Y} (self : CompressionSchemeWithInfo X Y C) {m : โ} (S : Fin m โ X ร Y),
โ(self.compress S).1 โ Set.range SCompressed set is small
โ {X : Type u} {Y : Type v} {C : ConceptClass X Y} (self : CompressionSchemeWithInfo X Y C) {m : โ} (S : Fin m โ X ร Y),
(self.compress S).1.card โค self.kernelSizeA labeled compression scheme with finite side information. This is the object proved to exist by Moran-Yehudayoff (2016, arXiv:1503.06960).
The current CompressionScheme is strictly stronger: it requires reconstruction from the compressed Finset alone (no side information). See Open_NoInfoCompressionStrengthening for that conjecture.
(X : Type u) โ (Y : Type v) โ ConceptClass X Y โ Type (max (max u (u_1 + 1)) v)
The side information type
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ CompressionSchemeWithInfo X Y C โ Type u_1Compression: extract โค kernelSize labeled examples + side information
{X : Type u} โ
{Y : Type v} โ
{C : ConceptClass X Y} โ
(self : CompressionSchemeWithInfo X Y C) โ {m : โ} โ (Fin m โ X ร Y) โ Finset (X ร Y) ร self.InfoSide information is finite
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ (self : CompressionSchemeWithInfo X Y C) โ Fintype self.InfoKernel size bound
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ CompressionSchemeWithInfo X Y C โ โReconstruction: produce hypothesis from compressed subset AND side information
{X : Type u} โ
{Y : Type v} โ {C : ConceptClass X Y} โ (self : CompressionSchemeWithInfo X Y C) โ Finset (X ร Y) โ self.Info โ X โ YTotal size of a compression scheme with side information: kernel size + number of side information states. (The paper uses k + logโ(|I|+1); we use the simpler k + |I| which is an upper bound and avoids importing Real.log.)
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ CompressionSchemeWithInfo X Y C โ โFix the hidden Info universe parameter of CompressionSchemeWithInfo to 0. This resolves the universe elaboration obstruction: Fin T โ Finset (Fin K) is Type 0, while CompressionSchemeWithInfo X Bool C with X : Type u infers Info : Type u. Pinning to .{u, 0, 0} allows Type 0 Info directly.
(X : Type u) โ (Y : Type) โ ConceptClass X Y โ Type (max (max u 1) 0)
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
Empirical error: average loss on a finite sample.
(X : Type u) โ (Y : Type v) โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ LossFunction Y โ โA concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
A probability mass function over a finite type. Named FinitePMF to avoid conflict with Mathlib's PMF.
(H : Type u_1) โ [Fintype H] โ Type u_1
{H : Type u_1} โ [inst : Fintype H] โ FinitePMF H โ H โ โ{H : Type u_1} โ [inst : Fintype H] โ FinitePMF H โ PMF HVC dimension of a finite Bool-valued family, computed via the set-system image boolFamilyToFinsetFamily and Mathlib's Finset.vcDim. Declared noncomputable because the underlying vcDim is.
{H : Type u_1} โ [Fintype H] โ [DecidableEq H] โ Finset (H โ Bool) โ โGrowth function (shattering coefficient): ฯ_C(m) = max_{|S|=m} |{c|_S : c โ C}|. For each m-element set S, counts the number of distinct restrictions of C to S, then takes the supremum over all such S.
(X : Type u) โ ConceptClass X Bool โ โ โ โ
Concrete side information for the MY construction: each of the T recovered blocks is represented by the set of kernel positions it uses.
โ โ โ โ Type
MWU config: weight vector with positivity proof.
(C : Type u_1) โ [Fintype C] โ Type u_1
Potential = sum of weights.
{C : Type u_1} โ [inst : Fintype C] โ MWUConfig C โ โNormalize config to PMF.
{C : Type u_1} โ [inst : Fintype C] โ [Nonempty C] โ MWUConfig C โ FinitePMF C{C : Type u_1} โ [inst : Fintype C] โ MWUConfig C โ C โ โA proper finite-support learner for a concept class C. This structure captures the existence of a bounded-support ERM with error at most 1/3 for any C-realizable finite distribution. CORRECTED: good_on_support returns Finset X (not Fin k โ X).
(X : Type u) โ ConceptClass X Bool โ Type u
{X : Type u} โ {C : ConceptClass X Bool} โ ProperFiniteSupportLearner X C โ {m : โ} โ (Fin m โ X ร Bool) โ X โ Bool{X : Type u} โ {C : ConceptClass X Bool} โ ProperFiniteSupportLearner X C โ โA set S โ X is shattered by concept class C if every labeling of S is realized by some concept in C.
(X : Type u) โ ConceptClass X Bool โ Finset X โ Prop
โ โ Type
True error (0-1 loss, realizable case): D-probability of disagreement. This is what PACLearnable's success event measures.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ ENNReal
True error in โ: for use in bounds involving subtraction/absolute value. COUNTER-1 of TrueError. The toReal bridge loses information when the measure is โค.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ โ
VC dimension of a concept class: the size of the largest shattered set. Returns โโ = WithTop โ.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
Per-point agreement test: for a fixed point x โ Y and concept c, maps hypothesis h to whether h(x) = c(x).
{X : Type u} โ (X โ Bool) โ X โ (HY : Finset (X โ Bool)) โ โฅHY โ BoolThe family of agreement tests over all points in Y.
{X : Type u} โ (X โ Bool) โ Finset X โ (HY : Finset (X โ Bool)) โ Finset (โฅHY โ Bool)Maps a finite family of Bool-valued functions to its image as a family of accepting sets. The set-system view is what Mathlib's Finset.Shatters and Finset.vcDim consume, so this is the entry point from the function-class view to the combinatorial VC machinery.
{H : Type u_1} โ [Fintype H] โ [DecidableEq H] โ Finset (H โ Bool) โ Finset (Finset H)Expected payoff of distribution p against column c in a Boolean game.
{R : Type u_1} โ {C : Type u_2} โ [inst : Fintype R] โ (R โ C โ Bool) โ FinitePMF R โ C โ โExpected value of a Bool-valued test under a finite distribution, via the indicator embedding if f h then 1 else 0. The central quantity of the finite-VC approximation layer: a TV bound on distributions translates to a uniform bound on test expectations via expectation_approx_of_tv.
{H : Type u_1} โ [inst : Fintype H] โ FinitePMF H โ (H โ Bool) โ โConvert Bool labels to ยฑ1 reals. true โฆ 1, false โฆ -1.
Bool โ โ
Bounded subsamples: all subsets of Y with cardinality โค s.
{X : Type u} โ Finset X โ โ โ Finset (Finset X)Decode labels from the kernel. This is exactly the current MY reconstruction convention in your file.
{X : Type u} โ [DecidableEq X] โ Finset (X ร Bool) โ X โ BoolDecode the X-coordinates of a block from kernel positions. This matches the current blockHyp shape.
{X : Type u} โ Finset (X ร Bool) โ {K : โ} โ Finset (Fin K) โ Finset XThe disagreement family: for each h โ C, the test y โฆ decide(h(y) โ c(y)) restricted to Y. Used for the VC approximation step in the proper learner proof.
{X : Type u} โ ConceptClass X Bool โ (X โ Bool) โ (Y : Finset X) โ Finset (โฅY โ Bool)Build FinitePMF from empirical frequencies of a finite sequence.
{ฮฑ : Type u_1} โ [inst : Fintype ฮฑ] โ [DecidableEq ฮฑ] โ {T : โ} โ 0 < T โ (Fin T โ ฮฑ) โ FinitePMF ฮฑEncode a witness set W as the set of kernel positions of the pairs (x, c x). The bound kernel.card โค K is fed into the encoding through the if branch, so the result has the same shape as the current compressCore code.
{X : Type u} โ [DecidableEq X] โ Finset (X ร Bool) โ (X โ Bool) โ (K : โ) โ Finset X โ Finset (Fin K){H : Type u_1} โ (H โ Bool) โ H โ โ โ BoolThe hypothesis envelope: the finite set of all possible learner outputs on bounded subsamples of Y, labeled by concept c.
{X : Type u} โ {C : ConceptClass X Bool} โ ProperFiniteSupportLearner X C โ (X โ Bool) โ Finset X โ Finset (X โ Bool)Build a labeled sample from a Finset of points and a concept.
{X : Type u} โ (X โ Bool) โ (Z : Finset X) โ Fin Z.card โ X ร Bool{H : Type u_1} โ Finset (H โ Bool) โ ConceptClass (H โ โ) BoolThe actual final closure helper. Packages the majority-vote construction. If decoded hypotheses agree with reference hypotheses on sample points, and majority of reference hypotheses agree with each label, then majority-vote reconstruction is correct.
{X : Type u} โ
{C : ConceptClass X Bool} โ
(T K : โ) โ
(compressCore : {m : โ} โ (Fin m โ X ร Bool) โ Finset (X ร Bool) ร IncidenceInfo T K) โ
(blockHyp : Finset (X ร Bool) โ IncidenceInfo T K โ Fin T โ X โ Bool) โ
(rowHyp :
{m : โ} โ (S : Fin m โ X ร Bool) โ (โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) โ Fin T โ X โ Bool) โ
0 < T โ
(โ {m : โ} (S : Fin m โ X ร Bool), (compressCore S).1.card โค K) โ
(โ {m : โ} (S : Fin m โ X ร Bool), โ(compressCore S).1 โ Set.range S) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m)
(t : Fin T),
blockHyp (compressCore S).1 (compressCore S).2 t (S i).1 = rowHyp S hreal t (S i).1) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m),
(โ t, if rowHyp S hreal t (S i).1 = (S i).2 then 1 else 0) / โT > 1 / 2) โ
CompressionSchemeWithInfo0 X Bool CThe MWU config after T steps.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ โ โ MWUConfig CCount how many rounds hit a fixed column, aligned to the recursion of mwuRun.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ (โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ โ โ C โ โInitial config: all weights = 1.
(C : Type u_1) โ [inst : Fintype C] โ MWUConfig C
The MWU row sequence after T steps.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ (T : โ) โ Fin T โ RMWU run: iterate T steps, returning final config and row sequence.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ
(T : โ) โ MWUConfig C ร (Fin T โ R)One MWU update step on weights.
{C : Type u_1} โ [inst : Fintype C] โ {R : Type u_2} โ (R โ C โ Bool) โ (ฮท : โ) โ ฮท < 1 โ MWUConfig C โ R โ MWUConfig CExtract the domain points from a labeled sample.
{X : Type u} โ {m : โ} โ (Fin m โ X ร Bool) โ Finset XWeighted error of hypothesis h vs concept c over a FinitePMF on Y.
{X : Type u} โ (Y : Finset X) โ FinitePMF โฅY โ (X โ Bool) โ (X โ Bool) โ โUniform PMF over a nonempty Fintype.
(C : Type u_1) โ [inst : Fintype C] โ [Nonempty C] โ FinitePMF C
The 0-1 loss for classification.
(Y : Type v) โ [DecidableEq Y] โ LossFunction Y
Littlestone characterization: C is online-learnable iff LittlestoneDim(C) < โ.
โ (X : Type) (C : ConceptClass X Bool), OnlineLearnable X Bool C โ LittlestoneDim X C < โค
โ (X : Type) (C : ConceptClass X Bool), OnlineLearnable X Bool C โ LittlestoneDim X C < โค
Forward direction: OnlineLearnable โ LittlestoneDim < โค
โ (X : Type) (C : ConceptClass X Bool), OnlineLearnable X Bool C โ LittlestoneDim X C < โค
Core adversary lemma.
โ {X : Type} (L : OnlineLearner X Bool) (s : L.State) {C : ConceptClass X Bool} {n : โ} (T : LTree X n),
LTree.isShattered C T โ Set.Nonempty C โ โ seq, โ c โ C, L.mistakesFrom s c seq = nBackward direction: LittlestoneDim < โค โ OnlineLearnable.
โ (X : Type) (C : ConceptClass X Bool), LittlestoneDim X C < โค โ OnlineLearnable X Bool C
SOA mistakes from a given state are bounded by the Ldim of the version space. This is the core M-Potential argument.
โ {X : Type} {C : ConceptClass X Bool},
โ c โ C,
โ (history : List (X ร Bool)),
(โ p โ history, c p.1 = p.2) โ
โ (d : โ),
LittlestoneDim X (versionSpace C history) โค โโd โ
LittlestoneDim X (versionSpace C history) < โค โ โ (seq : List X), (SOA X C).mistakesFrom history c seq โค dExtending history restricts the version space.
โ {X : Type} {C : ConceptClass X Bool} {history : List (X ร Bool)} {x : X} {y : Bool},
versionSpace C (history ++ [(x, y)]) โ versionSpace C historyTarget stays in version space.
โ {X : Type} {C : ConceptClass X Bool} {c : X โ Bool},
c โ C โ โ {history : List (X ร Bool)}, (โ p โ history, c p.1 = p.2) โ c โ versionSpace C historyOn an SOA mistake, the Ldim of the version space strictly decreases.
โ {X : Type} {C : ConceptClass X Bool} {history : List (X ร Bool)} {x : X} {c : X โ Bool},
c โ versionSpace C history โ
(SOA X C).predict history x โ c x โ
LittlestoneDim X (versionSpace C history) < โค โ
LittlestoneDim X (versionSpace C (history ++ [(x, c x)])) < LittlestoneDim X (versionSpace C history)SOA predicts the label whose side has higher Ldim.
โ {X : Type} (C : ConceptClass X Bool) (history : List (X ร Bool)) (x : X),
have V := versionSpace C history;
have b := (SOA X C).predict history x;
LittlestoneDim X {c | c โ V โง c x = b} โฅ LittlestoneDim X {c | c โ V โง c x = !b}When Ldim(V) = 0 (โโ0), all concepts in V agree on every point. Key lemma for M-VersionSpaceCollapse.
โ {X : Type} {V : ConceptClass X Bool},
LittlestoneDim X V = โ0 โ Set.Nonempty V โ โ (x : X) (cโ cโ : X โ Bool), cโ โ V โ cโ โ V โ cโ x = cโ xBuild a tree of depth k+1 from shattered subtrees on both sides. Parametrized over b : Bool so we don't need to case-split in the caller.
โ {X : Type} {V : ConceptClass X Bool} {x : X} {k : โ} {b : Bool},
(โ Tb, LTree.isShattered {c | c โ V โง c x = b} Tb) โ
(โ Tnb, LTree.isShattered {c | c โ V โง c x = !b} Tnb) โ
(โ c โ V, c x = b) โ (โ c โ V, c x = !b) โ LittlestoneDim X V โฅ โโ(k + 1)From Ldim โฅ d, extract a shattered tree of depth exactly d.
โ {X : Type} {C : ConceptClass X Bool} {d : โ}, LittlestoneDim X C โฅ โโd โ โ T, LTree.isShattered C TA truncated tree is shattered if the original is.
โ {X : Type} {C : ConceptClass X Bool} {m : โ} (T : LTree X m),
LTree.isShattered C T โ โ {n : โ} (h : n โค m), LTree.isShattered C (LTree.truncate h T)Helper: shattering implies the concept class is nonempty.
โ {X : Type} {C : ConceptClass X Bool} {n : โ} (T : LTree X n), LTree.isShattered C T โ Set.Nonempty CIn WithBot (WithTop โ), a < โโ(n+1) โ a โค โโn. Reusable lattice fact.
โ {a : WithBot (WithTop โ)} {n : โ}, a < โโ(n + 1) โ a โค โโnSOA mistakesFrom cons: unfold one step using the interface.
โ (X : Type) (C : ConceptClass X Bool) (history : List (X ร Bool)) (c : X โ Bool) (x : X) (xs : List X),
(SOA X C).mistakesFrom history c (x :: xs) =
(if (SOA X C).predict history x โ c x then 1 else 0) + (SOA X C).mistakesFrom (history ++ [(x, c x)]) c xsShattering is upward-monotone in the concept class.
โ {X : Type} {n : โ} (T : LTree X n) {C C' : ConceptClass X Bool},
C โ C' โ LTree.isShattered C T โ LTree.isShattered C' TRelate mistakesFrom to the original mistakes function.
โ {X : Type} (L : OnlineLearner X Bool) (c : X โ Bool) (seq : List X), L.mistakesFrom L.init c seq = L.mistakes c seqSOA init state is empty history.
โ (X : Type) (C : ConceptClass X Bool), (SOA X C).init = []
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
A concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
A complete binary Littlestone tree of depth n.
Type โ โ โ Type
Path-wise shattering for complete trees. Path B: leaf case requires C.Nonempty (NAโโ).
{X : Type} โ {n : โ} โ ConceptClass X Bool โ LTree X n โ PropTruncate a complete tree to a smaller depth.
{X : Type} โ {n m : โ} โ n โค m โ LTree X m โ LTree X nLittlestone dimension: the maximum depth of a complete shattered tree. Path B: returns WithBot (WithTop โ) so Ldim(โ ) = โฅ (NAโโ).
(X : Type) โ ConceptClass X Bool โ WithBot (WithTop โ)
Mistake-bounded learning: the learner makes at most M mistakes on ANY sequence. No distribution assumption. Characterized by Littlestone dimension.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ ConceptClass X Y โ โ โ Prop
Online learnable: there exists a finite mistake bound.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ ConceptClass X Y โ Prop
An online learner: receives instances one at a time, makes predictions sequentially.
Type u โ Type v โ Type (max (max 1 u) v)
Internal state type
{X : Type u} โ {Y : Type v} โ OnlineLearner X Y โ TypeInitial state
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.StateHelper: run an online learner on a sequence, counting mistakes.
{X : Type u} โ {Y : Type v} โ [DecidableEq Y] โ OnlineLearner X Y โ Concept X Y โ List X โ โ{X : Type u} โ {Y : Type v} โ [DecidableEq Y] โ (L : OnlineLearner X Y) โ Concept X Y โ L.State โ List X โ โ โ โCount mistakes starting from state s.
{X : Type} โ (L : OnlineLearner X Bool) โ L.State โ (X โ Bool) โ List X โ โPredict: given current state and new instance, output a prediction
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.State โ X โ YUpdate: given current state, instance, and revealed true label, update state
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.State โ X โ Y โ self.StateThe Standard Optimal Algorithm (SOA).
(X : Type) โ ConceptClass X Bool โ OnlineLearner X Bool
Version space after observing a history.
{X : Type} โ ConceptClass X Bool โ List (X ร Bool) โ ConceptClass X BoolThe optimal mistake bound equals the Littlestone dimension (for nonempty C). Path B: OptimalMistakeBound : WithTop โ, LittlestoneDim : WithBot (WithTop โ). For nonempty C, LittlestoneDim โฅ 0, so the coercion โ(OptimalMistakeBound) works.
โ (X : Type) (C : ConceptClass X Bool), Set.Nonempty C โ โ(OptimalMistakeBound X C) = LittlestoneDim X C
โ (X : Type) (C : ConceptClass X Bool), Set.Nonempty C โ โ(OptimalMistakeBound X C) = LittlestoneDim X C
Backward direction: LittlestoneDim < โค โ OnlineLearnable.
โ (X : Type) (C : ConceptClass X Bool), LittlestoneDim X C < โค โ OnlineLearnable X Bool C
SOA mistakes from a given state are bounded by the Ldim of the version space. This is the core M-Potential argument.
โ {X : Type} {C : ConceptClass X Bool},
โ c โ C,
โ (history : List (X ร Bool)),
(โ p โ history, c p.1 = p.2) โ
โ (d : โ),
LittlestoneDim X (versionSpace C history) โค โโd โ
LittlestoneDim X (versionSpace C history) < โค โ โ (seq : List X), (SOA X C).mistakesFrom history c seq โค dExtending history restricts the version space.
โ {X : Type} {C : ConceptClass X Bool} {history : List (X ร Bool)} {x : X} {y : Bool},
versionSpace C (history ++ [(x, y)]) โ versionSpace C historyTarget stays in version space.
โ {X : Type} {C : ConceptClass X Bool} {c : X โ Bool},
c โ C โ โ {history : List (X ร Bool)}, (โ p โ history, c p.1 = p.2) โ c โ versionSpace C historyOn an SOA mistake, the Ldim of the version space strictly decreases.
โ {X : Type} {C : ConceptClass X Bool} {history : List (X ร Bool)} {x : X} {c : X โ Bool},
c โ versionSpace C history โ
(SOA X C).predict history x โ c x โ
LittlestoneDim X (versionSpace C history) < โค โ
LittlestoneDim X (versionSpace C (history ++ [(x, c x)])) < LittlestoneDim X (versionSpace C history)SOA predicts the label whose side has higher Ldim.
โ {X : Type} (C : ConceptClass X Bool) (history : List (X ร Bool)) (x : X),
have V := versionSpace C history;
have b := (SOA X C).predict history x;
LittlestoneDim X {c | c โ V โง c x = b} โฅ LittlestoneDim X {c | c โ V โง c x = !b}When Ldim(V) = 0 (โโ0), all concepts in V agree on every point. Key lemma for M-VersionSpaceCollapse.
โ {X : Type} {V : ConceptClass X Bool},
LittlestoneDim X V = โ0 โ Set.Nonempty V โ โ (x : X) (cโ cโ : X โ Bool), cโ โ V โ cโ โ V โ cโ x = cโ xBuild a tree of depth k+1 from shattered subtrees on both sides. Parametrized over b : Bool so we don't need to case-split in the caller.
โ {X : Type} {V : ConceptClass X Bool} {x : X} {k : โ} {b : Bool},
(โ Tb, LTree.isShattered {c | c โ V โง c x = b} Tb) โ
(โ Tnb, LTree.isShattered {c | c โ V โง c x = !b} Tnb) โ
(โ c โ V, c x = b) โ (โ c โ V, c x = !b) โ LittlestoneDim X V โฅ โโ(k + 1)From Ldim โฅ d, extract a shattered tree of depth exactly d.
โ {X : Type} {C : ConceptClass X Bool} {d : โ}, LittlestoneDim X C โฅ โโd โ โ T, LTree.isShattered C TA truncated tree is shattered if the original is.
โ {X : Type} {C : ConceptClass X Bool} {m : โ} (T : LTree X m),
LTree.isShattered C T โ โ {n : โ} (h : n โค m), LTree.isShattered C (LTree.truncate h T)In WithBot (WithTop โ), a < โโ(n+1) โ a โค โโn. Reusable lattice fact.
โ {a : WithBot (WithTop โ)} {n : โ}, a < โโ(n + 1) โ a โค โโnSOA mistakesFrom cons: unfold one step using the interface.
โ (X : Type) (C : ConceptClass X Bool) (history : List (X ร Bool)) (c : X โ Bool) (x : X) (xs : List X),
(SOA X C).mistakesFrom history c (x :: xs) =
(if (SOA X C).predict history x โ c x then 1 else 0) + (SOA X C).mistakesFrom (history ++ [(x, c x)]) c xsShattering is upward-monotone in the concept class.
โ {X : Type} {n : โ} (T : LTree X n) {C C' : ConceptClass X Bool},
C โ C' โ LTree.isShattered C T โ LTree.isShattered C' TAdversary lower bound (M-InfSup reusable primitive): If a tree of depth n is shattered by C, then any mistake-bounded learner must allow at least n mistakes. This is the "inf โฅ sup" half of minimax.
โ {X : Type} {C : ConceptClass X Bool} {n : โ} (T : LTree X n),
LTree.isShattered C T โ โ {M : โ}, MistakeBounded X Bool C M โ n โค MRelate mistakesFrom to the original mistakes function.
โ {X : Type} (L : OnlineLearner X Bool) (c : X โ Bool) (seq : List X), L.mistakesFrom L.init c seq = L.mistakes c seqCore adversary lemma.
โ {X : Type} (L : OnlineLearner X Bool) (s : L.State) {C : ConceptClass X Bool} {n : โ} (T : LTree X n),
LTree.isShattered C T โ Set.Nonempty C โ โ seq, โ c โ C, L.mistakesFrom s c seq = nHelper: shattering implies the concept class is nonempty.
โ {X : Type} {C : ConceptClass X Bool} {n : โ} (T : LTree X n), LTree.isShattered C T โ Set.Nonempty CSOA init state is empty history.
โ (X : Type) (C : ConceptClass X Bool), (SOA X C).init = []
Set.Nonempty C
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
A concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
A complete binary Littlestone tree of depth n.
Type โ โ โ Type
Path-wise shattering for complete trees. Path B: leaf case requires C.Nonempty (NAโโ).
{X : Type} โ {n : โ} โ ConceptClass X Bool โ LTree X n โ PropTruncate a complete tree to a smaller depth.
{X : Type} โ {n m : โ} โ n โค m โ LTree X m โ LTree X nLittlestone dimension: the maximum depth of a complete shattered tree. Path B: returns WithBot (WithTop โ) so Ldim(โ ) = โฅ (NAโโ).
(X : Type) โ ConceptClass X Bool โ WithBot (WithTop โ)
Mistake-bounded learning: the learner makes at most M mistakes on ANY sequence. No distribution assumption. Characterized by Littlestone dimension.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ ConceptClass X Y โ โ โ Prop
Online learnable: there exists a finite mistake bound.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ ConceptClass X Y โ Prop
An online learner: receives instances one at a time, makes predictions sequentially.
Type u โ Type v โ Type (max (max 1 u) v)
Internal state type
{X : Type u} โ {Y : Type v} โ OnlineLearner X Y โ TypeInitial state
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.StateHelper: run an online learner on a sequence, counting mistakes.
{X : Type u} โ {Y : Type v} โ [DecidableEq Y] โ OnlineLearner X Y โ Concept X Y โ List X โ โ{X : Type u} โ {Y : Type v} โ [DecidableEq Y] โ (L : OnlineLearner X Y) โ Concept X Y โ L.State โ List X โ โ โ โCount mistakes starting from state s.
{X : Type} โ (L : OnlineLearner X Bool) โ L.State โ (X โ Bool) โ List X โ โPredict: given current state and new instance, output a prediction
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.State โ X โ YUpdate: given current state, instance, and revealed true label, update state
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.State โ X โ Y โ self.StateMistake bound: minimum worst-case mistakes for online learning of C.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
The Standard Optimal Algorithm (SOA).
(X : Type) โ ConceptClass X Bool โ OnlineLearner X Bool
Version space after observing a history.
{X : Type} โ ConceptClass X Bool โ List (X ร Bool) โ ConceptClass X BoolUniversal learnable โ PAC learnable. Proof sketch: UniversalLearnable gives learner L with rate โ 0 and Pr[error โค rate(m)] โฅ 2/3. Two components: 1. Event containment: rate(m) < ฮต โน {error โค rate(m)} โ {error โค ฮต} (monotonicity). 2. Confidence boosting: 2/3 โ 1-ฮด via median-of-means (ฮโโ, sorry'd in boost_two_thirds_to_pac). Routes through boost_two_thirds_to_pac which encapsulates the Chernoff-based boosting.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) [MeasurableHypotheses X C], (โ (L : BatchLearner X Bool), LearnEvalMeasurable L) โ UniversalLearnable X C โ PACLearnable X C
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) [MeasurableHypotheses X C], (โ (L : BatchLearner X Bool), LearnEvalMeasurable L) โ UniversalLearnable X C โ PACLearnable X C
Boosting lemma: given a learner with success probability โฅ 2/3 under D^m, construct a boosted learner with success probability โฅ 1-ฮด for any ฮด > 0. Standard technique: run L independently k times on independent samples of size mโ, take majority vote.
Construction:
ฮโโ: sorry โ the full measure-theoretic proof requires: (a) block_extract : (Fin (kn) โ X) โ Fin k โ (Fin n โ X) (b) iIndepFun for block extractions under product measure D^(kn) (c) chebyshev_majority_bound for i.i.d. Bernoulli(โฅ2/3) events (d) block extraction marginal = D^n (e) majority vote D-error analysis via union bound
None of this infrastructure currently exists in the codebase. The sorry is A4-compliant (the conclusion PACLearnable X C is non-trivially-true: it requires genuine concentration + majority analysis) and A5-compliant (the proof strategy is structurally complete, only infrastructure is missing).
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) [MeasurableHypotheses X C] (L : BatchLearner X Bool)
[MeasurableBatchLearner X L] (rate : โ โ โ),
(โ ฮต > 0, โ mโ, โ m โฅ mโ, rate m < ฮต) โ
(โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ
โ c โ C,
โ (m : โ),
(MeasureTheory.Measure.pi fun x => D)
{xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal (rate m)} โฅ
ENNReal.ofReal (2 / 3)) โ
PACLearnable X CT5: The goodBlockEvents are independent under the product measure.
โ {X : Type u} [inst : MeasurableSpace X] (L : BatchLearner X Bool) [MeasurableBatchLearner X L]
(D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D] (c : Concept X Bool),
Measurable c โ
โ (rate : โ โ โ) (k n : โ),
ProbabilityTheory.iIndepSet (goodBlockEventโ L D c rate k n) (MeasureTheory.Measure.pi fun x => D)Block extractions are independent under the product measure. Key infrastructure for boosting (D4) and probability amplification.
โ {X : Type u_1} [inst : MeasurableSpace X] (k m : โ) (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D],
ProbabilityTheory.iIndepFun (fun j ฯ => block_extract k m ฯ j) (MeasureTheory.Measure.pi fun x => D)T4: Each block's good event has probability โฅ 2/3, transported from the base learner guarantee.
โ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) (L : BatchLearner X Bool)
[MeasurableBatchLearner X L] (rate : โ โ โ),
(โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ
โ c โ C,
โ (m : โ),
(MeasureTheory.Measure.pi fun x => D)
{xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal (rate m)} โฅ
ENNReal.ofReal (2 / 3)) โ
โ (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D],
โ c โ C,
Measurable c โ
โ (k n : โ) (j : Fin k),
(MeasureTheory.Measure.pi fun x => D) (goodBlockEventโ L D c rate k n j) โฅ ENNReal.ofReal (2 / 3)T0: Shared helper โ measurability of the "good training set" event for a single block.
โ {X : Type u} [inst : MeasurableSpace X] (L : BatchLearner X Bool) [MeasurableBatchLearner X L]
(D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D] (c : Concept X Bool),
Measurable c โ
โ (rate : โ โ โ) (n : โ),
MeasurableSet {xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal (rate n)}T3: Block extraction pushforward of product measure equals product measure.
โ {X : Type u} [inst : MeasurableSpace X] (k n : โ) (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(j : Fin k),
MeasureTheory.Measure.map (fun ฯ => block_extract k n ฯ j) (MeasureTheory.Measure.pi fun x => D) =
MeasureTheory.Measure.pi fun x => DT2: goodBlockEvent is measurable for each block index j.
โ {X : Type u} [inst : MeasurableSpace X] (L : BatchLearner X Bool) [MeasurableBatchLearner X L]
(D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D] (c : Concept X Bool),
Measurable c โ โ (rate : โ โ โ) (k n : โ) (j : Fin k), MeasurableSet (goodBlockEventโ L D c rate k n j)Block extraction is measurable: extracting block j from a pi-type is measurable.
โ {X : Type u_1} [inst : MeasurableSpace X] (k m : โ) (j : Fin k), Measurable fun ฯ => block_extract k m ฯ jT6: Chebyshev concentration for 7/12 threshold โ when k โฅ 36/ฮด independent events each have probability โฅ 2/3, the fraction exceeding 7/12 is โฅ 1-ฮด.
โ {ฮฉ : Type u_1} [inst : MeasurableSpace ฮฉ] {ฮผ : MeasureTheory.Measure ฮฉ} [MeasureTheory.IsProbabilityMeasure ฮผ] {k : โ}
{ฮด : โ},
0 < ฮด โ
36 / ฮด โค โk โ
โ (events : Fin k โ Set ฮฉ),
(โ (j : Fin k), MeasurableSet (events j)) โ
ProbabilityTheory.iIndepSet (fun j => events j) ฮผ โ
(โ (j : Fin k), ฮผ (events j) โฅ ENNReal.ofReal (2 / 3)) โ
ฮผ {ฯ | 7 * k โค 12 * {j | ฯ โ events j}.card} โฅ ENNReal.ofReal (1 - ฮด)T8: If โฅ 7/12 of blocks are good, the boosted hypothesis has D-error โค 7ยทmax(rate(n),0).
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(c : Concept X Bool),
Measurable c โ
โ (L : BatchLearner X Bool) [MeasurableBatchLearner X L] (rate : โ โ โ) (k n : โ) (ฯ : Fin (k * n) โ X),
0 < k โ
7 * k โค 12 * {j | ฯ โ goodBlockEventโ L D c rate k n j}.card โ
D
{x |
(boosted_majorityโ k fun j =>
L.learn (fun i => (block_extract k n ฯ j i, c (block_extract k n ฯ j i))) x) โ
c x} โค
ENNReal.ofReal (7 * max (rate n) 0)T7: If โฅ 7/12 of the hypotheses have D-error โค ฯ, majority vote has D-error โค 7ฯ.
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D] {k : โ},
0 < k โ
โ (c : Concept X Bool),
Measurable c โ
โ (hs : Fin k โ Concept X Bool),
(โ (j : Fin k), Measurable (hs j)) โ
โ (good : Finset (Fin k)),
7 * k โค 12 * good.card โ
โ (ฯ : โ),
0 โค ฯ โ
(โ j โ good, D {x | hs j x โ c x} โค ENNReal.ofReal ฯ) โ
D {x | (boosted_majorityโ k fun j => hs j x) โ c x} โค ENNReal.ofReal (7 * ฯ)T1: A learner with joint measurability gives measurable hypotheses for fixed training data.
โ {X : Type u} [inst : MeasurableSpace X] (L : BatchLearner X Bool) [MeasurableBatchLearner X L] {m : โ}
(S : Fin m โ X ร Bool), Measurable (L.learn S)Joint measurability of the evaluation map
โ {X : Type u} {inst : MeasurableSpace X} {L : BatchLearner X Bool} [self : MeasurableBatchLearner X L] (m : โ),
Measurable fun p => L.learn p.1 p.2โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableHypotheses X C],
โ h โ C, Measurable hโ (L : BatchLearner X Bool), LearnEvalMeasurable L
UniversalLearnable X C
A batch learner (PAC paradigm): takes a finite sample, returns a hypothesis.
Type u โ Type v โ Type (max u v)
The learning algorithm: given a sample, produce a hypothesis
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ {m : โ} โ (Fin m โ X ร Y) โ Concept X YA concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
A concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
Joint measurability of a batch learner's evaluation map.
{X : Type u} โ [MeasurableSpace X] โ BatchLearner X Bool โ PropA batch learner whose evaluation map is jointly measurable.
The condition: for each sample size m, the map (S, x) โฆ L.learn S x from (Fin m โ X ร Bool) ร X to Bool is Measurable.
This is the minimal regularity that makes the PAC success event {S | D{x | L.learn(S)(x) โ c(x)} โค ฮต} a MeasurableSet (via measurable_measure_prod_mk_left).
Equivalent to LearnEvalMeasurable (Separation.lean) and AdviceEvalMeasurable (Extended.lean) for the non-advice case.
(X : Type u) โ [MeasurableSpace X] โ BatchLearner X Bool โ Prop
Every concept in C is a measurable function. Krapp-Wirth precondition: ฮ(h) โ ฮฃ_Z for all h โ H.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
PAC (Probably Approximately Correct) learning. The central definition of computational learning theory.
Sample space: Fin m โ X with i.i.d. product measure D^m. Labels: derived deterministically from target concept c (realizable case). Error: D-probability of disagreement between learner output and c.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Universal learning: distribution-free convergence rates. Strictly stronger than PAC.
Sample space: Fin m โ X with i.i.d. product measure D^m (matching PACLearnable). Labels: derived deterministically from target concept c (realizable case). The rate function converges to 0, and for every m, with probability โฅ 2/3 over D^m, the learner's error is at most rate(m).
ฮโโ fix: changed from existential Dm to Measure.pi (CNAโโ definition repair).
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Extract block j from a flat array of k*m elements, using finProdFinEquiv.
{ฮฑ : Type u_1} โ (k m : โ) โ (Fin (k * m) โ ฮฑ) โ Fin k โ Fin m โ ฮฑMajority vote: returns true iff strictly more than half the votes are true.
(k : โ) โ (Fin k โ Bool) โ Bool
The event that block j produces a hypothesis with D-error โค rate(n).
{X : Type u} โ
[inst : MeasurableSpace X] โ
BatchLearner X Bool โ MeasureTheory.Measure X โ Concept X Bool โ (โ โ โ) โ (k n : โ) โ Fin k โ Set (Fin (k * n) โ X)The KrappโWirth measurable-target separation holds unconditionally: the analytic non-Borel witness discharges the hypothesis of analytic_nonborel_set_gives_measTarget_separation.
KrappWirthSeparationMeasTarget
KrappWirthSeparationMeasTarget
Main separation theorem. Given any analytic non-Borel set A โ โ, the concept class obtained by parameterising singletonConcept (plus zeroConcept) over A is a concrete witness that WellBehavedVCMeasTarget is strictly weaker than the Krapp-Wirth Borel condition. The class is constructed as Set.range e for an evaluation map e : Bool ร ฮฒ โ Concept โ Bool built from a Polish parameterisation of A; post-construction, Set.range e equals singletonClassOn (Set.range g) where g realises A.
The class satisfies:
MeasurableHypotheses: every individual hypothesis is Borel (singletonClassOn_measurable).WellBehavedVCMeasTarget: the bad event is analytic (planarWitnessEvent_analytic lifted via singleton_badEvent_eq_preimage_planar), hence NullMeasurableSet by the Choquet bridge.KrappWirthWellBehaved: the bad event is not Borel (singleton_badEvent_not_measurable).The separation is realised by passing through the standard Borel space โ as the parameter space; the construction reuses no problem-specific fact beyond the existence of an analytic non-Borel subset of โ (Souslin's classical result), supplied in exists_measTarget_separation. The witness shows that the measurable-target variant proved in this kernel is a genuine improvement over the existing literature, not a restatement.
โ (A : Set โ), MeasureTheory.AnalyticSet A โ ยฌMeasurableSet A โ KrappWirthSeparationMeasTarget
For A non-Borel, the singleton bad event is non-Borel. Combine singleton_badEvent_eq_preimage_planar with planarWitnessEvent_not_measurable: the preimage of a non-Borel set under a measurable surjection cannot itself be Borel.
โ (A : Set โ), ยฌMeasurableSet A โ ยฌMeasurableSet (singletonBadEvent A)
The singleton bad event equals samplePair1ToPlane โปยน' planarWitnessEvent. The set equality that transports both analyticity and non-Borelness from the planar witness to the learning-theoretic bad event.
โ (A : Set โ), singletonBadEvent A = samplePair1ToPlane โปยน' planarWitnessEvent A
For A non-Borel, planarWitnessEvent A is non-Borel. The proof picks some a โ A and shows the vertical section y โฆ (a, y) pulls the planar event back to A itself: if the planar event were Borel, its preimage under this measurable map would be Borel too, contradicting the hypothesis on A.
โ (A : Set โ), ยฌMeasurableSet A โ ยฌMeasurableSet (planarWitnessEvent A)
โ {X : Type u} [inst : MeasurableSpace X] (h c : Concept X Bool) (m : โ) (p : (Fin m โ X) ร (Fin m โ X)),
oneSidedGhostGap h c m p โ ghostGapGrid mโ {X : Type u} [MeasurableSpace X] (h : Concept X Bool) {m : โ} (S : Fin m โ X ร Bool),
EmpiricalError X Bool h S (zeroOneLoss Bool) โ empErrGrid mClass-level corollary: every Borel-parameterized concept class with a measurable evaluation map satisfies WellBehavedVCMeasTarget. Composes borel_param_nullMeasurableSet_bad_event over all measurable targets c. The measurable-target variant of WellBehavedVC is what the kernel actually proves; the unrestricted variant remains open and is the subject of the Borel-analytic separation witness in Theorem/BorelAnalyticSeparation.lean.
โ {X : Type u} [inst : MeasurableSpace X] [inst_1 : TopologicalSpace X] [PolishSpace X] [BorelSpace X] {ฮ : Type u_1}
[inst_4 : MeasurableSpace ฮ] [StandardBorelSpace ฮ] (e : ฮ โ Concept X Bool),
(Measurable fun p => e p.1 p.2) โ WellBehavedVCMeasTarget X (Set.range e)Positive bridge. If a concept class is parameterized by a Borel measurable map ฮ โ Concept X from a standard Borel space ฮ, then the symmetrization bad event is analytic, hence NullMeasurableSet. The bad event is a Suslin projection of a Borel witness set (the projection along ฮ of {(ฮธ, p) | gap(eval ฮธ, p) โฅ ฮต / 2}), and projections of Borel sets are analytic by definition. This is the entry point through which Borel parameterization implies the regularity required by the fundamental theorem.
โ {X : Type u} [inst : MeasurableSpace X] [inst_1 : TopologicalSpace X] [PolishSpace X] [BorelSpace X] {ฮ : Type u_1}
[inst_4 : MeasurableSpace ฮ] [StandardBorelSpace ฮ] (e : ฮ โ Concept X Bool),
(Measurable fun p => e p.1 p.2) โ
โ (c : Concept X Bool),
Measurable c โ
โ (m : โ) (ฮต : โ) (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D],
MeasureTheory.NullMeasurableSet (paramBadEvent e c m ฮต) (GhostPairMeasure D m)The bad event (projection of witness set) is analytic. Projection of a Borel set from a StandardBorelSpace is analytic (Suslin). This is the key step: existential quantification over parameters produces an analytic (ฮฃโยน) set, which may not be Borel.
โ {X : Type u} [inst : TopologicalSpace X] [inst_1 : MeasurableSpace X] [BorelSpace X] [PolishSpace X] {ฮ : Type u_1}
[inst_4 : MeasurableSpace ฮ] [StandardBorelSpace ฮ] (e : ฮ โ Concept X Bool),
(Measurable fun p => e p.1 p.2) โ
โ (c : Concept X Bool), Measurable c โ โ (m : โ) (ฮต : โ), MeasureTheory.AnalyticSet (paramBadEvent e c m ฮต)The witness set {(ฮธ, p) | ghost-gap โฅ ฮต/2} is MeasurableSet when the evaluation map e and target c are measurable. This is the Borel half of the Borel-analytic bridge.
โ {X : Type u} [inst : MeasurableSpace X] {ฮ : Type u_1} [inst_1 : MeasurableSpace ฮ] (e : ฮ โ Concept X Bool),
(Measurable fun p => e p.1 p.2) โ
โ (c : Concept X Bool), Measurable c โ โ (m : โ) (ฮต : โ), MeasurableSet (paramWitnessSet e c m ฮต)Analytic subsets of the ghost sample space (Fin m โ X) ร (Fin m โ X) are NullMeasurableSet under the product probability measure. A specialisation of analyticSet_nullMeasurableSet from PureMath/AnalyticMeasurability.lean to the type the symmetrization argument actually consumes.
โ {X : Type u} [inst : TopologicalSpace X] [inst_1 : MeasurableSpace X] [BorelSpace X] [PolishSpace X] {m : โ}
{s : Set ((Fin m โ X) ร (Fin m โ X))},
MeasureTheory.AnalyticSet s โ
โ (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D],
MeasureTheory.NullMeasurableSet s (GhostPairMeasure D m)Analytic sets are null-measurable for finite Borel measures on Polish spaces. FLT-facing alias of the ZPM-canonical MeasureTheory.AnalyticSet.nullMeasurableSet.
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
{ฮผ : MeasureTheory.Measure ฮฑ} [MeasureTheory.IsFiniteMeasure ฮผ] {s : Set ฮฑ},
MeasureTheory.AnalyticSet s โ MeasureTheory.NullMeasurableSet s ฮผAnalytic sets are null-measurable. For any finite Borel measure on a Polish space, every analytic set is NullMeasurableSet. Proof inner- approximates the analytic set by compacts, takes the union of approximators (a Borel set), and shows the difference is contained in a Borel null set.
This is the abstract bridge that the entire Borel-analytic measurability layer of downstream learning-theory kernels rests on.
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
{ฮผ : MeasureTheory.Measure ฮฑ} [MeasureTheory.IsFiniteMeasure ฮผ] {s : Set ฮฑ},
MeasureTheory.AnalyticSet s โ MeasureTheory.NullMeasurableSet s ฮผInner regularity for analytic sets: any analytic subset of a Polish space can be approximated from inside by a compact subset in measure, by any slack ฮต > 0. Specialisation of Choquet capacitability (Kechris 30.13) to the measure-as-capacity instance.
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
{ฮผ : MeasureTheory.Measure ฮฑ} [MeasureTheory.IsFiniteMeasure ฮผ] {s : Set ฮฑ},
MeasureTheory.AnalyticSet s โ โ ฮต > 0, โ K, IsCompact K โง K โ s โง ฮผ.real s < ฮผ.real K + ฮตFor analytic sets, compact capacity equals measure.
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
{ฮผ : MeasureTheory.Measure ฮฑ} [MeasureTheory.IsFiniteMeasure ฮผ] {s : Set ฮฑ},
MeasureTheory.AnalyticSet s โ MeasureTheory.compactCap ฮผ s = ฮผ sโ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] (ฮผ : MeasureTheory.Measure ฮฑ) (s : Set ฮฑ),
MeasureTheory.compactCap ฮผ s = โจ K, โจ (_ : IsCompact K), โจ (_ : K โ s), ฮผ KEvery finite Borel measure on a Polish space is a Choquet capacity.
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
(ฮผ : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsFiniteMeasure ฮผ], MeasureTheory.IsChoquetCapacity fun s => ฮผ sChoquet capacitability. For analytic sets, capacity equals the supremum over compact subsets (Kechris 30.13).
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
{cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ
โ {s : Set ฮฑ}, MeasureTheory.AnalyticSet s โ cap s = โจ K, โจ (_ : IsCompact K), โจ (_ : K โ s), cap Kโ (N : โ โ โ) (n : โ), Monotone fun k => Cyl N n โฉ {g | g (n + 1) โค k}The intersection of closures of cylinder images equals the compact image. Key lemma for the capacitability proof: uses truncation and sequential compactness.
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [PolishSpace ฮฑ] {f : (โ โ โ) โ ฮฑ},
Continuous f โ โ (N : โ โ โ), โ n, closure (f '' Cyl N n) = f '' Bnd Nโ (N g : โ โ โ), choquetTruncate N g โ Bnd N
โ (N : โ โ โ) (n : โ), โ g โ Cyl N n, โ i โค n, choquetTruncate N g i = g i
โ (N : โ โ โ), IsCompact (Bnd N)
โ (N : โ โ โ) (n : โ), Bnd N โ Cyl N n
โ (N : โ โ โ) (n : โ), Cyl N n = โ k, Cyl N n โฉ {g | g (n + 1) โค k}โ (N : โ โ โ) (n k : โ), Cyl N n โฉ {g | g (n + 1) โค k} = Cyl (Function.update N (n + 1) k) (n + 1)โ (N N' : โ โ โ) (n : โ), (โ i โค n, N i = N' i) โ Cyl N n = Cyl N' n
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] {cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ โ {s t : Set ฮฑ}, s โ t โ cap s โค cap tโ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] {cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ โ (f : โ โ Set ฮฑ), Monotone f โ cap (โ n, f n) = โจ n, cap (f n)โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] {cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ
โ (f : โ โ Set ฮฑ), Antitone f โ (โ (n : โ), IsClosed (f n)) โ cap (โ n, f n) = โจ
n, cap (f n)โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : KrappWirthWellBehaved X C], KrappWirthV X CAn analytic non-Borel subset of โ: the image of the Baire-space witness under the continuous injection. Analyticity transfers along the continuous image; non-Borelness transfers back along the injective preimage.
โ A, MeasureTheory.AnalyticSet A โง ยฌMeasurableSet A
An analytic non-Borel subset of Baire space: the projection of the diagonal witness.
โ A, MeasureTheory.AnalyticSet A โง ยฌMeasurableSet A
A closed set with a non-Borel projection. Diagonalize the universal closed set of (โ โ โ) ร (โ โ โ): were the projection Borel, its complement would be analytic, hence the projection of a closed set, hence a section of the universal set โ and evaluating that section at its own parameter is contradictory.
โ D, IsClosed D โง ยฌMeasurableSet {x | โ y, (x, y) โ D}A universal closed set. Every second-countable space X carries a closed subset of X ร (โ โ โ) whose sections run through all closed subsets of X: enumerate a countable basis together with โ
, and let the parameter select which basis elements to exclude.
โ (X : Type u_1) [inst : TopologicalSpace X] [SecondCountableTopology X],
โ S, IsClosed S โง โ (C : Set X), IsClosed C โ โ y, {x | (x, y) โ S} = CFunction.Injective MeasureTheory.embedBaireReal
Function.Injective MeasureTheory.baireMarkerBits
โ (x : โ โ โ), StrictMono (MeasureTheory.baireMarkers x)
Continuous MeasureTheory.embedBaireReal
Continuous (Cardinal.cantorFunction (1 / 3))
Continuous MeasureTheory.baireMarkerBits
Bounded functions set: {g : โ โ โ | โ i, g i โค N i}.
(โ โ โ) โ Set (โ โ โ)
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
Cylinder set: {g : โ โ โ | โ i โค n, g i โค N i}.
(โ โ โ) โ โ โ Set (โ โ โ)
Empirical error: average loss on a finite sample.
(X : Type u) โ (Y : Type v) โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ LossFunction Y โ โA concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
{X : Type u} โ
[inst : MeasurableSpace X] โ MeasureTheory.Measure X โ (m : โ) โ MeasureTheory.Measure ((Fin m โ X) ร (Fin m โ X))Ghost sample pairs: two independent samples of size m.
Type u โ โ โ Type u
The ghost sample space at sample size m = 1: (Fin 1 โ โ) ร (Fin 1 โ โ). The smallest sample size at which the singleton-class obstruction is already visible.
Type
OPEN QUESTION (measurable-target version): Does WellBehavedVCMeasTarget separate from KrappWirthWellBehaved? The Borel-analytic bridge (BorelAnalyticBridge.lean) closes this.
Prop
V-measurability (one-sided): the ghost gap sup map is measurable.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Krapp-Wirth well-behavedness: measurable hypotheses + V + U. Extends MeasurableHypotheses (L1). Strictly stronger than MeasurableConceptClass (our condition).
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Every concept in C is a measurable function. Krapp-Wirth precondition: ฮ(h) โ ฮฃ_Z for all h โ H.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Bundled record of the three Choquet capacity axioms: monotonicity, sequential continuity from below along increasing unions, and sequential continuity from above along decreasing intersections of closed sets. The third axiom distinguishes a capacity from a general outer measure.
{ฮฑ : Type u_1} โ [TopologicalSpace ฮฑ] โ (Set ฮฑ โ ENNReal) โ PropThe marker bits: the indicator stream of the marker set.
(โ โ โ) โ โ โ Bool
The marker sequence of x : โ โ โ: the strictly increasing sequence n + 1 + โ_{k โค n} x k, whose successive gaps encode x.
(โ โ โ) โ โ โ โ
Compact capacity of a set s relative to a measure ฮผ: the supremum of ฮผ K over compact subsets K โ s. The inner-regularity functional whose equality with ฮผ s characterises measurability for analytic sets.
{ฮฑ : Type u_1} โ [TopologicalSpace ฮฑ] โ [inst : MeasurableSpace ฮฑ] โ MeasureTheory.Measure ฮฑ โ Set ฮฑ โ ENNRealThe embedding of Baire space into โ: marker bits into the base-3 expansion.
(โ โ โ) โ โ
WellBehavedVC restricted to measurable targets. This is the correct target for the Borel-analytic positive bridge: Borel parameterization โ analytic bad event โ NullMeasurableSet, but only when c is measurable (so the ghost-gap map is measurable).
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Truncation: replace g i by min (g i) (N i) to bring any g into the bounded set.
(โ โ โ) โ (โ โ โ) โ โ โ โ
โ โ Finset โ
โ โ Finset โ
{X : Type u} โ [MeasurableSpace X] โ ConceptClass X Bool โ Concept X Bool โ (m : โ) โ (Fin m โ X) ร (Fin m โ X) โ โ{X : Type u} โ [MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ (m : โ) โ (Fin m โ X) ร (Fin m โ X) โ โThe bad event in sample space: projection of the witness set. Existential over the parameter: {p | โ ฮธ, gap(ฮธ, p) โฅ ฮต/2}. This is analytic when the witness set is Borel (Theorem B).
{X : Type u} โ
[MeasurableSpace X] โ
{ฮ : Type u_1} โ [MeasurableSpace ฮ] โ (ฮ โ Concept X Bool) โ Concept X Bool โ (m : โ) โ โ โ Set (GhostPairs X m)The witness set in parameter ร sample space: {(ฮธ, p) | EmpErr(h_ฮธ, ghost, c) - EmpErr(h_ฮธ, train, c) โฅ ฮต/2}. This is Borel when e and c are measurable (Theorem A).
{X : Type u} โ
[MeasurableSpace X] โ
{ฮ : Type u_1} โ
[MeasurableSpace ฮ] โ (ฮ โ Concept X Bool) โ Concept X Bool โ (m : โ) โ โ โ Set (ฮ ร GhostPairs X m)The planar witness {(x, y) โ โ ร โ | y โ A โง x โ y}. For A analytic non-Borel, this set is itself analytic non-Borel. The geometric core of the separation: the learning-theoretic bad event below is a measurable preimage of this planar set.
Set โ โ Set (โ ร โ)
The projection GhostPairs1 โ โ ร โ, p โฆ (p.1 0, p.2 0). Surjective and measurable; non-Borelness of a target set transfers to non-Borelness of its preimage under a measurable surjection.
GhostPairs1 โ โ ร โ
The symmetrization bad event for the singleton class at sample size m = 1, target concept zeroConcept, and threshold 1/2. Equals the preimage of planarWitnessEvent under samplePair1ToPlane (see singleton_badEvent_eq_preimage_planar), and inherits both analyticity and non-Borelness from the planar set when A is analytic non-Borel.
Set โ โ Set GhostPairs1
The singleton class over A โ โ: {zeroConcept} โช {singletonConcept a | a โ A}. The zeroConcept disjunct is the target concept against which the symmetrization bad event is measured. For A analytic non-Borel, this is the witness used to separate WellBehavedVCMeasTarget from the Krapp-Wirth Borel condition.
Set โ โ ConceptClass โ Bool
The point indicator singletonConcept a x = (x = a). Each singletonConcept a is itself Borel measurable; non-Borelness in the singleton-class witness comes from quantifying over a โ A for A analytic non-Borel, not from any individual concept.
โ โ Concept โ Bool
The constantly false concept. The base hypothesis of the singleton class, serving both as the target concept and as the zeroConcept disjunct of singletonClassOn.
Concept โ Bool
The 0-1 loss for classification.
(Y : Type v) โ [DecidableEq Y] โ LossFunction Y
The Ambainis composition lock, at every depth: the product partition of the (d+1)-fold Ambainis iterate โ the construction giving the multiplicative upper bound on the partition number โ is not the leaf partition of any deterministic protocol, for every depth d. Index 0 is the base partition itself (a kernel-checked eight-rectangle monochromatic partition of the base game); index 1 is the 64-rectangle object of the depth-two frontier.
โ (d : โ),
ยฌโ t,
KWLock.Realizes (KWLock.onesOf (KWLock.Fd 4 KWLock.fA (d + 1))) (KWLock.zerosOf (KWLock.Fd 4 KWLock.fA (d + 1)))
(KWLock.towerP 4 KWLock.fA KWLock.pA d) tโ (d : โ),
ยฌโ t,
KWLock.Realizes (KWLock.onesOf (KWLock.Fd 4 KWLock.fA (d + 1))) (KWLock.zerosOf (KWLock.Fd 4 KWLock.fA (d + 1)))
(KWLock.towerP 4 KWLock.fA KWLock.pA d) tThe composition lock. Under the base hypotheses, the product partition of the (d+1)-fold iterate is not the leaf partition of any deterministic protocol, for every depth.
โ (k : โ) (f : (Fin k โ Bool) โ Bool) (Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))),
KWLock.TowerHyp k f Pโ โ
โ (d : โ),
ยฌโ t,
KWLock.Realizes (KWLock.onesOf (KWLock.Fd k f (d + 1))) (KWLock.zerosOf (KWLock.Fd k f (d + 1)))
(KWLock.towerP k f Pโ d) tEverything transfers up the tower: at every depth the product partition is a valid, monochromatic partition of the iterate's game, with both overlap graphs connected and two distinct members.
โ (k : โ) (f : (Fin k โ Bool) โ Bool) (Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))),
KWLock.TowerHyp k f Pโ โ
โ (d : โ),
KWLock.IsPartition (KWLock.onesOf (KWLock.Fd k f (d + 1))) (KWLock.zerosOf (KWLock.Fd k f (d + 1)))
(KWLock.towerP k f Pโ d) โง
KWLock.MonoValid (KWLock.vval k d) (KWLock.vval k d) (KWLock.towerP k f Pโ d) โง
KWLock.RowConnected (KWLock.towerP k f Pโ d) โง
KWLock.ColConnected (KWLock.towerP k f Pโ d) โง
โ r โ KWLock.towerP k f Pโ d, โ s โ KWLock.towerP k f Pโ d, r โ sPointwise-equal Boolean functions have equal zeros.
โ {ฮฑ : Type} [inst : Fintype ฮฑ] {hโ hโ : ฮฑ โ Bool}, (โ (x : ฮฑ), hโ x = hโ x) โ KWLock.zerosOf hโ = KWLock.zerosOf hโPointwise-equal Boolean functions have equal ones.
โ {ฮฑ : Type} [inst : Fintype ฮฑ] {hโ hโ : ฮฑ โ Bool}, (โ (x : ฮฑ), hโ x = hโ x) โ KWLock.onesOf hโ = KWLock.onesOf hโThe connectivity transfer, rows. Outer connectivity plus block connectivity (both graphs of every block partition) yields row connectivity of the product.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) (f : (Fin k โ Bool) โ Bool)
{Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.IsPartition (KWLock.onesOf f) (KWLock.zerosOf f) Po โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
KWLock.RowConnected Po โ
(โ (i : Fin k), KWLock.RowConnected (Q i)) โ
(โ (i : Fin k), KWLock.ColConnected (Q i)) โ KWLock.RowConnected (KWLock.compPartition g Po Q)The hop. Outer rectangles sharing a row yield touching composed rectangles: directly across distinct blocks, and through a common block rectangle when the blocks coincide (in which case the shared pattern forces equal orientations).
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
โ {Ro So : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)},
Ro โ Po โ
So โ Po โ
KWLock.RowTouch Ro So โ
โ {q : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)},
q โ Q Ro.label โ โ p โ Q So.label, KWLock.RowTouch (KWLock.compRect g Ro q) (KWLock.compRect g So p)Within one outer rectangle, composed rectangles over linked block rectangles are linked: the cluster inherits the block partition's row graph in the standard orientation and its column graph in the reversed one.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) (f : (Fin k โ Bool) โ Bool)
{Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.IsPartition (KWLock.onesOf f) (KWLock.zerosOf f) Po โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
(โ (i : Fin k), KWLock.RowConnected (Q i)) โ
(โ (i : Fin k), KWLock.ColConnected (Q i)) โ
โ {Ro : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)},
Ro โ Po โ
โ {q q' : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)},
q โ Q Ro.label โ
q' โ Q Ro.label โ
Relation.ReflTransGen (KWLock.TStep KWLock.RowTouch (KWLock.compPartition g Po Q))
(KWLock.compRect g Ro q) (KWLock.compRect g Ro q')Same-outer composed rectangles sharing a side value share a row.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
โ {Ro : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)},
Ro โ Po โ
Ro.rows.Nonempty โ
โ {q q' : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)},
q โ Q Ro.label โ
โ {vโ : V Ro.label},
vโ โ KWLock.rowSideSel Ro.orient q โ
vโ โ KWLock.rowSideSel Ro.orient q' โ
KWLock.RowTouch (KWLock.compRect g Ro q) (KWLock.compRect g Ro q')Labels transfer. The composed partition is monochromatic at its physical coordinates, under the composed valuation.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))} (vฮบ : (i : Fin k) โ V i โ ฮบ i โ Bool),
(โ (i : Fin k), KWLock.MonoValid (vฮบ i) (vฮบ i) (Q i)) โ
KWLock.MonoValid (fun w s => vฮบ s.fst (w s.fst) s.snd) (fun w s => vฮบ s.fst (w s.fst) s.snd)
(KWLock.compPartition g Po Q)Distinctness transfers: children of two distinct outer rectangles are distinct, because their cells are nonempty and outer-disjoint.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) (f : (Fin k โ Bool) โ Bool)
{Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.IsPartition (KWLock.onesOf f) (KWLock.zerosOf f) Po โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
(โ r โ Po, โ s โ Po, r โ s) โ โ r โ KWLock.compPartition g Po Q, โ s โ KWLock.compPartition g Po Q, r โ sExtract pairwise cell-disjointness for two distinct members.
โ {X Y ฮน : Type} {P : List (KWLock.KRect X Y ฮน)},
List.Pairwise (fun r s => โ (x : X) (y : Y), ยฌ(r.Cell x y โง s.Cell x y)) P โ
โ {r s : KWLock.KRect X Y ฮน}, r โ P โ s โ P โ r โ s โ โ (x : X) (y : Y), ยฌ(r.Cell x y โง s.Cell x y)Validity transfers. The product of a valid labeled outer partition with valid block partitions is a valid partition of the composed game.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) (f : (Fin k โ Bool) โ Bool)
{Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
KWLock.IsPartition (KWLock.onesOf f) (KWLock.zerosOf f) Po โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.IsPartition (KWLock.onesOf (KWLock.compFun g f)) (KWLock.zerosOf (KWLock.compFun g f))
(KWLock.compPartition g Po Q)A value on the row side of a block rectangle evaluates, under the block function, to the orientation โ provided the block partition is contained in its game.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ Fintype (V i)] {ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool)
{a : Fin k} {q : KWLock.KRect (V a) (V a) (ฮบ a)},
q.rows โ KWLock.onesOf (g a) โ
q.cols โ KWLock.zerosOf (g a) โ โ (o : Bool) {v : V a}, v โ KWLock.rowSideSel o q โ g a v = oRow sides of members of a valid block partition are nonempty.
โ {k : โ} {V ฮบ : Fin k โ Type} {a : Fin k} {q : KWLock.KRect (V a) (V a) (ฮบ a)},
q.rows.Nonempty โง q.cols.Nonempty โ โ (o : Bool), (KWLock.rowSideSel o q).NonemptyPairwise cell-disjointness transfers to the product: a shared composed cell projects to a shared outer cell (across outer rectangles) or a shared block cell in the orientation's order (within one outer rectangle).
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
List.Pairwise (fun r s => โ (x y : Fin k โ Bool), ยฌ(r.Cell x y โง s.Cell x y)) Po โ
(โ (i : Fin k), List.Pairwise (fun r s => โ (x y : V i), ยฌ(r.Cell x y โง s.Cell x y)) (Q i)) โ
List.Pairwise (fun r s => โ (x y : (i : Fin k) โ V i), ยฌ(r.Cell x y โง s.Cell x y)) (KWLock.compPartition g Po Q)โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Ro : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)}
{q : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)} {w : (i : Fin k) โ V i},
w โ (KWLock.compRect g Ro q).rows โ KWLock.blockPat g w โ Ro.rows โง w Ro.label โ KWLock.rowSideSel Ro.orient qโ {X Y ฮน : Type} {RX : Finset X} {CY : Finset Y} {P : List (KWLock.KRect X Y ฮน)},
KWLock.IsPartition RX CY P โ List.Pairwise (fun r s => โ (x : X) (y : Y), ยฌ(r.Cell x y โง s.Cell x y)) PThe connectivity transfer, columns.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) (f : (Fin k โ Bool) โ Bool)
{Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.IsPartition (KWLock.onesOf f) (KWLock.zerosOf f) Po โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
KWLock.ColConnected Po โ
(โ (i : Fin k), KWLock.RowConnected (Q i)) โ
(โ (i : Fin k), KWLock.ColConnected (Q i)) โ KWLock.ColConnected (KWLock.compPartition g Po Q)The column mirror of cross_step_row.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
โ {Ro So : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)},
Ro โ Po โ
So โ Po โ
KWLock.ColTouch Ro So โ
โ {q : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)},
q โ Q Ro.label โ โ p โ Q So.label, KWLock.ColTouch (KWLock.compRect g Ro q) (KWLock.compRect g So p)Build a composed point with a prescribed block pattern and prescribed values at two distinct blocks.
โ {k : โ} {V : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool),
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
โ (u : Fin k โ Bool) {aโ aโ : Fin k},
aโ โ aโ โ
โ (vโ : V aโ) (vโ : V aโ),
g aโ vโ = u aโ โ g aโ vโ = u aโ โ โ w, KWLock.blockPat g w = u โง w aโ = vโ โง w aโ = vโColumn sides of members of a valid block partition are nonempty.
โ {k : โ} {V ฮบ : Fin k โ Type} {a : Fin k} {q : KWLock.KRect (V a) (V a) (ฮบ a)},
q.rows.Nonempty โง q.cols.Nonempty โ โ (o : Bool), (KWLock.colSideSel o q).NonemptyBlock partitions of a game with a surjective block function are nonempty lists.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ Fintype (V i)] {ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool)
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ โ (i : Fin k), Q i โ []A valid nonempty-sided partition of a nonempty product is a nonempty list.
โ {X Y ฮน : Type} {RX : Finset X} {CY : Finset Y} {P : List (KWLock.KRect X Y ฮน)},
KWLock.IsPartition RX CY P โ RX.Nonempty โ CY.Nonempty โ P โ []โ {X Y ฮน : Type} {RX : Finset X} {CY : Finset Y} {P : List (KWLock.KRect X Y ฮน)},
KWLock.IsPartition RX CY P โ โ x โ RX, โ y โ CY, โ r โ P, r.Cell x yThe column mirror of cluster_linked_row.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) (f : (Fin k โ Bool) โ Bool)
{Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.IsPartition (KWLock.onesOf f) (KWLock.zerosOf f) Po โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
(โ (i : Fin k), KWLock.RowConnected (Q i)) โ
(โ (i : Fin k), KWLock.ColConnected (Q i)) โ
โ {Ro : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)},
Ro โ Po โ
โ {q q' : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)},
q โ Q Ro.label โ
q' โ Q Ro.label โ
Relation.ReflTransGen (KWLock.TStep KWLock.ColTouch (KWLock.compPartition g Po Q))
(KWLock.compRect g Ro q) (KWLock.compRect g Ro q')โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))}
{r : KWLock.KRect ((i : Fin k) โ V i) ((i : Fin k) โ V i) ((i : Fin k) ร ฮบ i)},
r โ KWLock.compPartition g Po Q โ โ Ro โ Po, โ q โ Q Ro.label, r = KWLock.compRect g Ro qAlong linkage from a member, every node is a member.
โ {X Y ฮน : Type} (T : KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop) {P : List (KWLock.KRect X Y ฮน)}
{r s : KWLock.KRect X Y ฮน}, r โ P โ Relation.ReflTransGen (KWLock.TStep T P) r s โ s โ PSame-outer composed rectangles sharing a side value share a column.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Po : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))}
{Q : (i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))},
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
KWLock.MonoValid (fun u i => u i) (fun u i => u i) Po โ
(โ (i : Fin k), KWLock.IsPartition (KWLock.onesOf (g i)) (KWLock.zerosOf (g i)) (Q i)) โ
โ {Ro : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)},
Ro โ Po โ
Ro.cols.Nonempty โ
โ {q q' : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)},
q โ Q Ro.label โ
โ {vโ : V Ro.label},
vโ โ KWLock.colSideSel Ro.orient q โ
vโ โ KWLock.colSideSel Ro.orient q' โ
KWLock.ColTouch (KWLock.compRect g Ro q) (KWLock.compRect g Ro q')โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ DecidableEq (V i)] [inst_1 : (i : Fin k) โ Fintype (V i)]
{ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) {Ro : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)}
{q : KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label)} {w : (i : Fin k) โ V i},
w โ (KWLock.compRect g Ro q).cols โ KWLock.blockPat g w โ Ro.cols โง w Ro.label โ KWLock.colSideSel Ro.orient qBuild a composed point with a prescribed block pattern and a prescribed value at one block, from surjectivity of the block functions.
โ {k : โ} {V : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool),
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ
โ (u : Fin k โ Bool) (a : Fin k) (v : V a), g a v = u a โ โ w, KWLock.blockPat g w = u โง w a = vA value on the column side of a block rectangle evaluates to the negated orientation.
โ {k : โ} {V : Fin k โ Type} [inst : (i : Fin k) โ Fintype (V i)] {ฮบ : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool)
{a : Fin k} {q : KWLock.KRect (V a) (V a) (ฮบ a)},
q.rows โ KWLock.onesOf (g a) โ
q.cols โ KWLock.zerosOf (g a) โ โ (o : Bool) {v : V a}, v โ KWLock.colSideSel o q โ g a v = !oโ {ฮฑ : Type} [inst : Fintype ฮฑ] {h : ฮฑ โ Bool} {x : ฮฑ}, x โ KWLock.zerosOf h โ h x = falseโ {ฮฑ : Type} [inst : Fintype ฮฑ] {h : ฮฑ โ Bool} {x : ฮฑ}, x โ KWLock.onesOf h โ h x = trueโ {X Y ฮน : Type} {RX : Finset X} {CY : Finset Y} {P : List (KWLock.KRect X Y ฮน)},
KWLock.IsPartition RX CY P โ โ r โ P, r.rows โ RX โง r.cols โ CYโ {k : โ} {f : (Fin k โ Bool) โ Bool} {Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))},
KWLock.TowerHyp k f Pโ โ KWLock.RowConnected Pโโ {k : โ} {f : (Fin k โ Bool) โ Bool} {Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))},
KWLock.TowerHyp k f Pโ โ KWLock.IsPartition (KWLock.onesOf f) (KWLock.zerosOf f) Pโโ {k : โ} {f : (Fin k โ Bool) โ Bool} {Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))},
KWLock.TowerHyp k f Pโ โ โ (b : Bool), โ u, f u = bโ {k : โ} {f : (Fin k โ Bool) โ Bool} {Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))},
KWLock.TowerHyp k f Pโ โ KWLock.MonoValid (fun u i => u i) (fun u i => u i) Pโโ {k : โ} {f : (Fin k โ Bool) โ Bool} {Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))},
KWLock.TowerHyp k f Pโ โ โ r โ Pโ, โ s โ Pโ, r โ sโ {k : โ} {f : (Fin k โ Bool) โ Bool} {Pโ : List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k))},
KWLock.TowerHyp k f Pโ โ KWLock.ColConnected PโThe iterate is surjective at every level.
โ (k : โ) (f : (Fin k โ Bool) โ Bool), (โ (b : Bool), โ u, f u = b) โ โ (d : โ) (b : Bool), โ w, KWLock.Fd k f (d + 1) w = b
Surjectivity composes.
โ {k : โ} {V : Fin k โ Type} (g : (i : Fin k) โ V i โ Bool) (f : (Fin k โ Bool) โ Bool),
(โ (b : Bool), โ u, f u = b) โ
(โ (i : Fin k) (b : Bool), โ v, g i v = b) โ โ (b : Bool), โ w, KWLock.compFun g f w = bThe atomicity obstruction. Any list of nonempty-sided rectangles with two distinct members whose row-overlap graph and column-overlap graph are both connected is not the leaf list of any deterministic protocol โ full partition validity is not needed. At the root split, every leaf's live side lies within one half; sharing a row (column) forces the same half; connectivity drags every rectangle to one half; yet both subtrees own a leaf.
โ {X Y ฮน : Type} [inst : DecidableEq X] [inst_1 : DecidableEq Y] {RX : Finset X} {CY : Finset Y}
{P : List (KWLock.KRect X Y ฮน)},
(โ r โ P, r.rows.Nonempty โง r.cols.Nonempty) โ
KWLock.RowConnected P โ KWLock.ColConnected P โ (โ r โ P, โ s โ P, r โ s) โ ยฌโ t, KWLock.Realizes RX CY P tLying inside a fixed half of a row split is invariant along the row-overlap graph, provided every member of P lies inside one half or the other: a shared row cannot be on both sides.
โ {X Y ฮน : Type} [inst : DecidableEq X] {P : List (KWLock.KRect X Y ฮน)} {RX A : Finset X},
(โ p โ P, p.rows โ RX โฉ A โจ p.rows โ RX \ A) โ
โ {p q : KWLock.KRect X Y ฮน},
Relation.ReflTransGen (KWLock.TStep KWLock.RowTouch P) p q โ p.rows โ RX โฉ A โ q.rows โ RX โฉ AThe column mirror of row_side_invariant.
โ {X Y ฮน : Type} [inst : DecidableEq Y] {P : List (KWLock.KRect X Y ฮน)} {CY B : Finset Y},
(โ p โ P, p.cols โ CY โฉ B โจ p.cols โ CY \ B) โ
โ {p q : KWLock.KRect X Y ฮน},
Relation.ReflTransGen (KWLock.TStep KWLock.ColTouch P) p q โ p.cols โ CY โฉ B โ q.cols โ CY โฉ Bโ {X Y ฮน : Type} (t : KWLock.Protocol X Y ฮน), t.leaves โ []Every leaf of a following protocol sits inside the current rectangle.
โ {X Y ฮน : Type} [inst : DecidableEq X] [inst_1 : DecidableEq Y] {RX : Finset X} {CY : Finset Y}
(t : KWLock.Protocol X Y ฮน), KWLock.Protocol.Follows RX CY t โ โ r โ t.leaves, r.rows โ RX โง r.cols โ CYโ {X Y ฮน : Type} {RX : Finset X} {CY : Finset Y} {P : List (KWLock.KRect X Y ฮน)},
KWLock.IsPartition RX CY P โ โ r โ P, r.rows.Nonempty โง r.cols.NonemptyThe Ambainis base satisfies every tower hypothesis.
KWLock.TowerHyp 4 KWLock.fA KWLock.pA
The row-overlap graph is connected: an explicit spanning walk.
KWLock.RowConnected KWLock.pA
โ {X Y ฮน : Type} {r s : KWLock.KRect X Y ฮน}, KWLock.RowTouch r s โ KWLock.RowTouch s rKernel re-check: every rectangle is monochromatic at its label, in its orientation.
KWLock.MonoValid (fun u i => u i) (fun u i => u i) KWLock.pA
Kernel re-check: the eight rectangles form a valid partition of the Ambainis game.
KWLock.IsPartition (KWLock.onesOf KWLock.fA) (KWLock.zerosOf KWLock.fA) KWLock.pA
Two distinct members.
โ r โ KWLock.pA, โ s โ KWLock.pA, r โ s
The column-overlap graph is connected: an explicit spanning walk.
KWLock.ColConnected KWLock.pA
A spanning walk โ inside P, visiting every member โ yields connectivity.
โ {X Y ฮน : Type} (T : KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop),
(โ {r s : KWLock.KRect X Y ฮน}, T r s โ T s r) โ
โ {P w : List (KWLock.KRect X Y ฮน)},
w โ [] โ KWLock.IsWalk T w โ (โ r โ w, r โ P) โ (โ r โ P, r โ w) โ KWLock.ConnectedVia T PThe head of a walk inside P links to every element of the walk.
โ {X Y ฮน : Type} (T : KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop) {P : List (KWLock.KRect X Y ฮน)}
{a : KWLock.KRect X Y ฮน} {l : List (KWLock.KRect X Y ฮน)},
KWLock.IsWalk T (a :: l) โ (โ r โ a :: l, r โ P) โ โ b โ a :: l, Relation.ReflTransGen (KWLock.TStep T P) a bFrom a member of P, every linked rectangle is a member, and the linkage reverses when the touch relation is symmetric.
โ {X Y ฮน : Type} (T : KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop),
(โ {r s : KWLock.KRect X Y ฮน}, T r s โ T s r) โ
โ {P : List (KWLock.KRect X Y ฮน)} {r s : KWLock.KRect X Y ฮน},
r โ P โ Relation.ReflTransGen (KWLock.TStep T P) r s โ s โ P โง Relation.ReflTransGen (KWLock.TStep T P) s rโ {X Y ฮน : Type} {r s : KWLock.KRect X Y ฮน}, KWLock.ColTouch r s โ KWLock.ColTouch s rThe base function is surjective.
โ (b : Bool), โ u, KWLock.fA u = b
Column connectivity.
{X Y ฮน : Type} โ List (KWLock.KRect X Y ฮน) โ PropTwo rectangles touch on columns when some column lies in both column sets.
{X Y ฮน : Type} โ KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ PropThe touch graph of P is connected.
{X Y ฮน : Type} โ (KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop) โ List (KWLock.KRect X Y ฮน) โ PropThe d-th iterate of the base function.
(k : โ) โ ((Fin k โ Bool) โ Bool) โ (d : โ) โ KWLock.Sp k d โ Bool
A valid partition of the product RX ร CY into nonempty rectangles: contained, pairwise cell-disjoint, covering.
{X Y ฮน : Type} โ Finset X โ Finset Y โ List (KWLock.KRect X Y ฮน) โ PropA walk: consecutive elements touch.
{X Y ฮน : Type} โ (KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop) โ List (KWLock.KRect X Y ฮน) โ PropA labeled rectangle: row and column sets, a coordinate label, and the rectangle's orientation (true when rows carry value true at the label). The atomicity results never read label/orient; the composition and validity results do.
Type โ Type โ Type โ Type
Cell membership.
{X Y ฮน : Type} โ KWLock.KRect X Y ฮน โ X โ Y โ Prop{X Y ฮน : Type} โ KWLock.KRect X Y ฮน โ Finset Y{X Y ฮน : Type} โ KWLock.KRect X Y ฮน โ ฮน{X Y ฮน : Type} โ KWLock.KRect X Y ฮน โ Bool{X Y ฮน : Type} โ KWLock.KRect X Y ฮน โ Finset XPhysical coordinates at depth d: paths of blocks ending in a base coordinate.
โ โ โ โ Type
Per-rectangle monochromaticity relative to valuations: rows carry the rectangle's orientation at its label, columns the negation.
{X Y ฮน : Type} โ (X โ ฮน โ Bool) โ (Y โ ฮน โ Bool) โ List (KWLock.KRect X Y ฮน) โ PropA deterministic protocol tree: the row player splits the current row set, the column player the current column set; leaves announce rectangles.
Type โ Type โ Type โ Type
The protocol respects the current rectangle (RX, CY): each split partitions the live side, and every leaf rectangle sits inside its branch's constraints.
{X Y ฮน : Type} โ [DecidableEq X] โ [DecidableEq Y] โ Finset X โ Finset Y โ KWLock.Protocol X Y ฮน โ Prop{X Y ฮน : Type} โ
{motive : KWLock.Protocol X Y ฮน โ Sort u} โ
(t : KWLock.Protocol X Y ฮน) โ
((t : KWLock.Protocol X Y ฮน) โ KWLock.Protocol.below t โ motive t) โ motive t ร' KWLock.Protocol.below tThe leaves of a protocol tree.
{X Y ฮน : Type} โ KWLock.Protocol X Y ฮน โ List (KWLock.KRect X Y ฮน)t realizes the partition P on (RX, CY): it follows the tree constraints and its leaves are exactly the members of P.
{X Y ฮน : Type} โ
[DecidableEq X] โ [DecidableEq Y] โ Finset X โ Finset Y โ List (KWLock.KRect X Y ฮน) โ KWLock.Protocol X Y ฮน โ PropRow connectivity.
{X Y ฮน : Type} โ List (KWLock.KRect X Y ฮน) โ PropTwo rectangles touch on rows when some row lies in both row sets.
{X Y ฮน : Type} โ KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ PropThe input space of the d-th iterate: Sp 1 is the base pattern space.
โ โ โ โ Type
One linkage step inside P: the target is a member and the rectangles touch.
{X Y ฮน : Type} โ
(KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop) โ
List (KWLock.KRect X Y ฮน) โ KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ PropThe base hypotheses of the tower: a valid, monochromatic, doubly-connected base partition with two distinct members, over a surjective base function.
(k : โ) โ ((Fin k โ Bool) โ Bool) โ List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)) โ Prop
The block pattern of a composed point: the vector of block outputs.
{k : โ} โ {V : Fin k โ Type} โ ((i : Fin k) โ V i โ Bool) โ ((i : Fin k) โ V i) โ Fin k โ BoolThe column side of a block rectangle, in a given orientation.
{k : โ} โ {V ฮบ : Fin k โ Type} โ Bool โ {a : Fin k} โ KWLock.KRect (V a) (V a) (ฮบ a) โ Finset (V a)The composed function.
{k : โ} โ {V : Fin k โ Type} โ ((i : Fin k) โ V i โ Bool) โ ((Fin k โ Bool) โ Bool) โ ((i : Fin k) โ V i) โ BoolThe product partition of the composed game.
{k : โ} โ
{V : Fin k โ Type} โ
[(i : Fin k) โ DecidableEq (V i)] โ
[(i : Fin k) โ Fintype (V i)] โ
{ฮบ : Fin k โ Type} โ
((i : Fin k) โ V i โ Bool) โ
List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)) โ
((i : Fin k) โ List (KWLock.KRect (V i) (V i) (ฮบ i))) โ
List (KWLock.KRect ((i : Fin k) โ V i) ((i : Fin k) โ V i) ((i : Fin k) ร ฮบ i))The composed rectangle: outer patterns from the outer rectangle, the active block constrained to the block rectangle's side in the outer orientation, other blocks free. The label is the physical coordinate; the orientation composes.
{k : โ} โ
{V : Fin k โ Type} โ
[(i : Fin k) โ DecidableEq (V i)] โ
[(i : Fin k) โ Fintype (V i)] โ
{ฮบ : Fin k โ Type} โ
((i : Fin k) โ V i โ Bool) โ
(Ro : KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)) โ
KWLock.KRect (V Ro.label) (V Ro.label) (ฮบ Ro.label) โ
KWLock.KRect ((i : Fin k) โ V i) ((i : Fin k) โ V i) ((i : Fin k) ร ฮบ i)The four-variable Ambainis function, in Ueno's ten-leaf form (truth table 0xd18b: exactly the monotone four-bit sequences evaluate to one).
(Fin 4 โ Bool) โ Bool
{X Y ฮน : Type} โ
[DecidableEq X] โ [DecidableEq Y] โ (r : KWLock.KRect X Y ฮน) โ (x : X) โ (y : Y) โ Decidable (r.Cell x y){X Y ฮน : Type} โ [DecidableEq Y] โ (r s : KWLock.KRect X Y ฮน) โ Decidable (KWLock.ColTouch r s){X Y ฮน : Type} โ [DecidableEq X] โ [DecidableEq Y] โ [DecidableEq ฮน] โ DecidableEq (KWLock.KRect X Y ฮน)(k d : โ) โ DecidableEq (KWLock.Sp k d)
{X Y ฮน : Type} โ
(T : KWLock.KRect X Y ฮน โ KWLock.KRect X Y ฮน โ Prop) โ
[(r s : KWLock.KRect X Y ฮน) โ Decidable (T r s)] โ (w : List (KWLock.KRect X Y ฮน)) โ Decidable (KWLock.IsWalk T w){X Y ฮน : Type} โ [DecidableEq X] โ (r s : KWLock.KRect X Y ฮน) โ Decidable (KWLock.RowTouch r s)(k d : โ) โ Fintype (KWLock.Sp k d)
The ones of a Boolean function, as a finite set.
{ฮฑ : Type} โ [Fintype ฮฑ] โ (ฮฑ โ Bool) โ Finset ฮฑThe witness partition.
List (KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4))
The eight-rectangle base partition, found by exact search and re-verified by the kernel below: eight monochromatic rectangles of eight cells each, in mixed orientations.
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
KWLock.KRect (Fin 4 โ Bool) (Fin 4 โ Bool) (Fin 4)
The row side of a block rectangle, in a given orientation (the transposed side for the reversed orientation).
{k : โ} โ {V ฮบ : Fin k โ Type} โ Bool โ {a : Fin k} โ KWLock.KRect (V a) (V a) (ฮบ a) โ Finset (V a)The tower of product partitions: depth 0 is the base partition; each next depth composes the base as the outer partition with the previous depth at every block.
(k : โ) โ
((Fin k โ Bool) โ Bool) โ
List (KWLock.KRect (Fin k โ Bool) (Fin k โ Bool) (Fin k)) โ
(d : โ) โ List (KWLock.KRect (KWLock.Sp k (d + 1)) (KWLock.Sp k (d + 1)) (KWLock.Lab k d))The valuation of an iterate point at a physical coordinate.
(k d : โ) โ KWLock.Sp k (d + 1) โ KWLock.Lab k d โ Bool
The zeros of a Boolean function, as a finite set.
{ฮฑ : Type} โ [Fintype ฮฑ] โ (ฮฑ โ Bool) โ Finset ฮฑIf every candidate's loss is measurable and a.e. valued in [a, b] under Q, every training-generated trajectory is admissible for F, and the resulting uniform-deviation failure event is measurable, then on an i.i.d. panel of size n drawn from Q, measured cumulative empirical-loss progress exceeds population-loss progress by more than 2 ยท deltaFiniteExperts (card ฮน) n (b - a) ฮด with probability at most ENNReal.ofReal ฮด.
โ {ฮฉtrain : Type u_1} {Z : Type u_2} {ฮน : Type u_3} [inst : MeasurableSpace ฮฉtrain] [inst_1 : MeasurableSpace Z]
[inst_2 : Fintype ฮน] [Nonempty ฮน] (ฮผtrain : AuditCP.AuditSampleLaw ฮฉtrain) [MeasureTheory.IsProbabilityMeasure ฮผtrain]
(Q : AuditCP.AuditSampleLaw Z) [MeasureTheory.IsProbabilityMeasure Q] (n : โ),
0 < n โ
โ (a b : โ),
a < b โ
โ (ฮด : โ),
0 < ฮด โ
ฮด โค 1 โ
โ (loss : ฮน โ Z โ โ),
(โ (i : ฮน), Measurable (loss i)) โ
(โ (i : ฮน), โแต (z : Z) โQ, loss i z โ Set.Icc a b) โ
โ (F : ฮฉtrain โ AuditCP.AuditEnvelope ฮน) (g : ฮฉtrain โ AuditCP.AuditTrajectory ฮน),
(โ (t : ฮฉtrain), AuditCP.Admissible (g t) (F t)) โ
MeasurableSet
{p |
ยฌAuditCP.UniformDev (F p.1) (fun i => AuditCP.empiricalLoss loss i p.2)
(fun i => AuditCP.populationLoss Q loss i)
(AuditCP.deltaFiniteExperts (Fintype.card ฮน) n (b - a) ฮด)} โ
โ (T : โ),
(MeasureTheory.Measure.prod ฮผtrain (MeasureTheory.Measure.pi fun x => Q))
{p |
AuditCP.cumCP (fun i => AuditCP.empiricalLoss loss i p.2) (g p.1) T >
AuditCP.cumCP (fun i => AuditCP.populationLoss Q loss i) (g p.1) T +
2 * AuditCP.deltaFiniteExperts (Fintype.card ฮน) n (b - a) ฮด} โค
ENNReal.ofReal ฮดโ {ฮฉtrain : Type u_1} {Z : Type u_2} {ฮน : Type u_3} [inst : MeasurableSpace ฮฉtrain] [inst_1 : MeasurableSpace Z]
[inst_2 : Fintype ฮน] [Nonempty ฮน] (ฮผtrain : AuditCP.AuditSampleLaw ฮฉtrain) [MeasureTheory.IsProbabilityMeasure ฮผtrain]
(Q : AuditCP.AuditSampleLaw Z) [MeasureTheory.IsProbabilityMeasure Q] (n : โ),
0 < n โ
โ (a b : โ),
a < b โ
โ (ฮด : โ),
0 < ฮด โ
ฮด โค 1 โ
โ (loss : ฮน โ Z โ โ),
(โ (i : ฮน), Measurable (loss i)) โ
(โ (i : ฮน), โแต (z : Z) โQ, loss i z โ Set.Icc a b) โ
โ (F : ฮฉtrain โ AuditCP.AuditEnvelope ฮน) (g : ฮฉtrain โ AuditCP.AuditTrajectory ฮน),
(โ (t : ฮฉtrain), AuditCP.Admissible (g t) (F t)) โ
MeasurableSet
{p |
ยฌAuditCP.UniformDev (F p.1) (fun i => AuditCP.empiricalLoss loss i p.2)
(fun i => AuditCP.populationLoss Q loss i)
(AuditCP.deltaFiniteExperts (Fintype.card ฮน) n (b - a) ฮด)} โ
โ (T : โ),
(MeasureTheory.Measure.prod ฮผtrain (MeasureTheory.Measure.pi fun x => Q))
{p |
AuditCP.cumCP (fun i => AuditCP.empiricalLoss loss i p.2) (g p.1) T >
AuditCP.cumCP (fun i => AuditCP.populationLoss Q loss i) (g p.1) T +
2 * AuditCP.deltaFiniteExperts (Fintype.card ฮน) n (b - a) ฮด} โค
ENNReal.ofReal ฮดUnder the hypotheses of finite_experts_iid_uniformDev, the ENNReal-valued probability that the panel fails uniform deviation is at most ENNReal.ofReal ฮด.
โ {Z : Type u_1} {ฮน : Type u_2} [inst : MeasurableSpace Z] [inst_1 : Fintype ฮน] [Nonempty ฮน]
(Q : AuditCP.AuditSampleLaw Z) [MeasureTheory.IsProbabilityMeasure Q] (n : โ),
0 < n โ
โ (a b : โ),
a < b โ
โ (ฮด : โ),
0 < ฮด โ
ฮด โค 1 โ
โ (loss : ฮน โ Z โ โ),
(โ (i : ฮน), Measurable (loss i)) โ
(โ (i : ฮน), โแต (z : Z) โQ, loss i z โ Set.Icc a b) โ
(MeasureTheory.Measure.pi fun x => Q)
{panel |
ยฌAuditCP.UniformDev Set.univ (fun i => AuditCP.empiricalLoss loss i panel)
(fun i => AuditCP.populationLoss Q loss i)
(AuditCP.deltaFiniteExperts (Fintype.card ฮน) n (b - a) ฮด)} โค
ENNReal.ofReal ฮดIf every candidate's loss is measurable and a.e. valued in [a, b] under Q, then on an i.i.d. panel of size n drawn from Q, with probability at least 1 โ ฮด the empirical loss empiricalLoss loss i panel and the population loss populationLoss Q loss i agree within deltaFiniteExperts (card ฮน) n (b - a) ฮด for every candidate i (hoeffding1963; cesabianchi2006).
โ {Z : Type u_1} {ฮน : Type u_2} [inst : MeasurableSpace Z] [inst_1 : Fintype ฮน] [Nonempty ฮน]
(Q : AuditCP.AuditSampleLaw Z) [MeasureTheory.IsProbabilityMeasure Q] (n : โ),
0 < n โ
โ (a b : โ),
a < b โ
โ (ฮด : โ),
0 < ฮด โ
ฮด โค 1 โ
โ (loss : ฮน โ Z โ โ),
(โ (i : ฮน), Measurable (loss i)) โ
(โ (i : ฮน), โแต (z : Z) โQ, loss i z โ Set.Icc a b) โ
1 - ฮด โค
(MeasureTheory.Measure.pi fun x => Q).real
{panel |
AuditCP.UniformDev Set.univ (fun i => AuditCP.empiricalLoss loss i panel)
(fun i => AuditCP.populationLoss Q loss i)
(AuditCP.deltaFiniteExperts (Fintype.card ฮน) n (b - a) ฮด)}If each Y i j is ฮผ-sub-Gaussian with proxy (L/2)ยฒ and, for each candidate i, independent across panel positions j, then with probability at least 1 โ ฮด the panel means (โโฑผ Y i j) / n all lie within deltaFiniteExperts (card ฮน) n L ฮด of 0 (hoeffding1963).
โ {ฮฉ : Type u_1} {ฮน : Type u_2} [inst : MeasurableSpace ฮฉ] [inst_1 : Fintype ฮน] [Nonempty ฮน]
(ฮผ : AuditCP.AuditSampleLaw ฮฉ) [MeasureTheory.IsProbabilityMeasure ฮผ] (n : โ),
0 < n โ
โ (L : โ),
0 < L โ
โ (ฮด : โ),
0 < ฮด โ
ฮด โค 1 โ
โ (Y : ฮน โ Fin n โ ฮฉ โ โ),
(โ (i : ฮน) (j : Fin n), ProbabilityTheory.HasSubgaussianMGF (Y i j) ((โLโโ / 2) ^ 2) ฮผ) โ
(โ (i : ฮน), ProbabilityTheory.iIndepFun (Y i) ฮผ) โ
1 - ฮด โค
MeasureTheory.Measure.real ฮผ
{ฯ |
AuditCP.UniformDev Set.univ (fun i => (โ j, Y i j ฯ) / โn) (fun x => 0)
(AuditCP.deltaFiniteExperts (Fintype.card ฮน) n L ฮด)}If every training-generated trajectory g t is admissible for F t, the uniform-deviation failure event is measurable, and its audit fiber has probability at most ฮด for ฮผtrain-almost every training history, then measured cumulative compression progress exceeds true progress by more than 2ฮ with probability at most ฮด under ฮผtrain.prod ฮผaudit.
โ {ฮฉtrain : Type u_1} {ฮฉaudit : Type u_2} {ฮน : Type u_3} [inst : MeasurableSpace ฮฉtrain]
[inst_1 : MeasurableSpace ฮฉaudit] (ฮผtrain : AuditCP.AuditSampleLaw ฮฉtrain) [MeasureTheory.IsProbabilityMeasure ฮผtrain]
(ฮผaudit : AuditCP.AuditSampleLaw ฮฉaudit) [MeasureTheory.IsProbabilityMeasure ฮผaudit]
(F : ฮฉtrain โ AuditCP.AuditEnvelope ฮน) (g : ฮฉtrain โ AuditCP.AuditTrajectory ฮน)
(Ehat : ฮฉaudit โ AuditCP.AuditPotential ฮน) (E : AuditCP.AuditPotential ฮน) (ฮ : โ) (ฮด : ENNReal),
(โ (t : ฮฉtrain), AuditCP.Admissible (g t) (F t)) โ
MeasurableSet {p | ยฌAuditCP.UniformDev (F p.1) (Ehat p.2) E ฮ} โ
(โแต (t : ฮฉtrain) โฮผtrain, ฮผaudit {a | ยฌAuditCP.UniformDev (F t) (Ehat a) E ฮ} โค ฮด) โ
โ (T : โ),
(MeasureTheory.Measure.prod ฮผtrain ฮผaudit)
{p | AuditCP.cumCP (Ehat p.2) (g p.1) T > AuditCP.cumCP E (g p.1) T + 2 * ฮ} โค
ฮดUnder UniformDev F Ehat E ฮ and Admissible g F, empirical cumulative compression progress is bounded by true cumulative progress plus 2ฮ at every horizon T: cumCP Ehat g T โค cumCP E g T + 2ฮ.
โ {ฮน : Type u_1} {F : AuditCP.AuditEnvelope ฮน} {Ehat E : AuditCP.AuditPotential ฮน} {ฮ : โ},
AuditCP.UniformDev F Ehat E ฮ โ
โ {g : AuditCP.AuditTrajectory ฮน},
AuditCP.Admissible g F โ โ (T : โ), AuditCP.cumCP Ehat g T โค AuditCP.cumCP E g T + 2 * ฮCumulative signed compression progress telescopes: cumCP E g T = E (g 0) - E (g T).
โ {ฮน : Type u_1} (E : AuditCP.AuditPotential ฮน) (g : AuditCP.AuditTrajectory ฮน) (T : โ),
AuditCP.cumCP E g T = E (g 0) - E (g T)If the event {p | Bad (ฮ p.1) p.2} is measurable and its audit fiber {a | Bad (ฮ t) a} has ฮผaudit-probability at most ฮด for ฮผtrain-almost every training history t, then its probability under the product law ฮผtrain.prod ฮผaudit is at most ฮด.
โ {ฮฉtrain : Type u_1} {ฮฉaudit : Type u_2} {ฮน : Type u_3} [inst : MeasurableSpace ฮฉtrain]
[inst_1 : MeasurableSpace ฮฉaudit] (ฮผtrain : AuditCP.AuditSampleLaw ฮฉtrain) [MeasureTheory.IsProbabilityMeasure ฮผtrain]
(ฮผaudit : AuditCP.AuditSampleLaw ฮฉaudit) [MeasureTheory.IsProbabilityMeasure ฮผaudit]
(ฮ : ฮฉtrain โ AuditCP.AuditEnvelope ฮน) (Bad : AuditCP.AuditEnvelope ฮน โ ฮฉaudit โ Prop) (ฮด : ENNReal),
MeasurableSet {p | Bad (ฮ p.1) p.2} โ
(โแต (t : ฮฉtrain) โฮผtrain, ฮผaudit {a | Bad (ฮ t) a} โค ฮด) โ
(MeasureTheory.Measure.prod ฮผtrain ฮผaudit) {p | Bad (ฮ p.1) p.2} โค ฮด0 < n
a < b
0 < ฮด
ฮด โค 1
โ (i : ฮน), Measurable (loss i)
โ (i : ฮน), โแต (z : Z) โQ, loss i z โ Set.Icc a b
โ (t : ฮฉtrain), AuditCP.Admissible (g t) (F t)
MeasurableSet
{p |
ยฌAuditCP.UniformDev (F p.1) (fun i => AuditCP.empiricalLoss loss i p.2) (fun i => AuditCP.populationLoss Q loss i)
(AuditCP.deltaFiniteExperts (Fintype.card ฮน) n (b - a) ฮด)}Every stage of the trajectory g lies in the envelope F.
{ฮน : Type u_1} โ AuditCP.AuditTrajectory ฮน โ AuditCP.AuditEnvelope ฮน โ PropB4 โ AuditEnvelope ฮน is the type of subsets Set ฮน of the index ฮน.
AuditCP.AuditIndex โ Type u
B2 โ AuditPotential ฮน is the type of real-valued readings ฮน โ โ of the index ฮน.
AuditCP.AuditIndex โ Type u
B7 โ AuditSampleLaw ฮฉ is MeasureTheory.Measure ฮฉ, a ฯ-additive probability law on the measurable space ฮฉ.
(ฮฉ : AuditCP.AuditIndex) โ [MeasurableSpace ฮฉ] โ Type u
B3 โ AuditTrajectory ฮน is the type of stage-indexed selections โ โ ฮน of the index ฮน.
AuditCP.AuditIndex โ Type u
The potentials Ehat and E agree to within ฮ at every point of the envelope F.
{ฮน : Type u_1} โ AuditCP.AuditEnvelope ฮน โ AuditCP.AuditPotential ฮน โ AuditCP.AuditPotential ฮน โ โ โ PropCumulative signed compression progress of a potential E read along a trajectory g over T stages, โ_{t < T} (E(g t) โ E(g(t+1))) (schmidhuber1991; schmidhuber2010).
{ฮน : Type u_1} โ AuditCP.AuditPotential ฮน โ AuditCP.AuditTrajectory ฮน โ โ โ โThe finite-experts uniform-deviation radius, Lยทโ(log(2N/ฮด)/(2n)), for N candidates, panel size n, loss range L, and failure probability ฮด (hoeffding1963).
โ โ โ โ โ โ โ โ โ
The empirical mean loss of candidate i on the panel, (โโฑผ loss i (panel j)) / n.
{Z : Type u_1} โ {ฮน : Type u_2} โ {n : โ} โ (ฮน โ Z โ โ) โ ฮน โ (Fin n โ Z) โ โThe population loss of candidate i under the law Q, โซ loss i z โQ.
{Z : Type u_1} โ {ฮน : Type u_2} โ [inst : MeasurableSpace Z] โ AuditCP.AuditSampleLaw Z โ (ฮน โ Z โ โ) โ ฮน โ โF2 โ given A nonempty, BlackwellEquivalent (learnerExperiment K) publicOnlyExperiment โ AuditSealed K (blackwell1953).
โ {S : Type uS} {A : Type uA} {Y : Type uY} [Nonempty A] (K : AuditCP.AuditChannel S A Y),
AuditCP.BlackwellEquivalent (AuditCP.learnerExperiment K) AuditCP.publicOnlyExperiment โ AuditCP.AuditSealed Kโ {S : Type uS} {A : Type uA} {Y : Type uY} [Nonempty A] (K : AuditCP.AuditChannel S A Y),
AuditCP.BlackwellEquivalent (AuditCP.learnerExperiment K) AuditCP.publicOnlyExperiment โ AuditCP.AuditSealed KF2a โ public information is Blackwell-below the complete learner view: BlackwellBelow publicOnlyExperiment (learnerExperiment K).
โ {S : Type uS} {A : Type uA} {Y : Type uY} (K : AuditCP.AuditChannel S A Y),
AuditCP.BlackwellBelow AuditCP.publicOnlyExperiment (AuditCP.learnerExperiment K)F2b โ given A nonempty, BlackwellBelow (learnerExperiment K) publicOnlyExperiment โ AuditSealed K (blackwell1953).
โ {S : Type uS} {A : Type uA} {Y : Type uY} [Nonempty A] (K : AuditCP.AuditChannel S A Y),
AuditCP.BlackwellBelow (AuditCP.learnerExperiment K) AuditCP.publicOnlyExperiment โ AuditCP.AuditSealed KPushing a PMF on Y through the tag y โฆ (s, y) is injective.
โ {S : Type uS} {Y : Type uY} (s : S), Function.Injective fun p => PMF.map (fun y => (s, y)) pF1 โ given A nonempty, AuditSealed K โ FactorsThroughPublic K.
โ {S : Type uS} {A : Type uA} {Y : Type uY} [Nonempty A] (K : AuditCP.AuditChannel S A Y),
AuditCP.AuditSealed K โ AuditCP.FactorsThroughPublic KB6 โ AuditChannel S A Y is AuditExperiment (S ร A) Y, an experiment whose parameter is split into a public coordinate S and an audit coordinate A.
AuditCP.AuditIndex โ AuditCP.AuditIndex โ AuditCP.AuditIndex โ Type (max u v w)
B5 โ AuditExperiment ฮ X is the type of parameter-indexed discrete observation laws ฮ โ PMF X.
AuditCP.AuditIndex โ AuditCP.AuditIndex โ Type (max u v)
The channel K is audit-sealed: its output law at fixed public state s does not depend on the audit state a.
{S : Type uS} โ {A : Type uA} โ {Y : Type uY} โ AuditCP.AuditChannel S A Y โ PropEโ is Blackwell-below Eโ when some garbling G of Eโ's output reproduces Eโ: โ G, โ ฮธ, Eโ ฮธ = (Eโ ฮธ).bind G (blackwell1953).
{ฮ : Type uฮ} โ {Xโ : Type uXโ} โ {Xโ : Type uXโ} โ AuditCP.AuditExperiment ฮ Xโ โ AuditCP.AuditExperiment ฮ Xโ โ PropEโ and Eโ are Blackwell-equivalent when each is Blackwell-below the other (blackwell1953).
{ฮ : Type uฮ} โ {Xโ : Type uXโ} โ {Xโ : Type uXโ} โ AuditCP.AuditExperiment ฮ Xโ โ AuditCP.AuditExperiment ฮ Xโ โ PropThe channel K factors through the public coordinate: โ H, โ s a, K (s, a) = H s.
{S : Type uS} โ {A : Type uA} โ {Y : Type uY} โ AuditCP.AuditChannel S A Y โ PropThe channel revealing the public state together with K's output, fun sa => (K sa).map (fun y => (sa.1, y)).
{S : Type uS} โ {A : Type uA} โ {Y : Type uY} โ AuditCP.AuditChannel S A Y โ AuditCP.AuditChannel S A (S ร Y)The channel revealing exactly the public state s and nothing else, fun (s, a) => PMF.pure s.
{S : Type uS} โ {A : Type uA} โ AuditCP.AuditChannel S A SF3 โ AuditSealed K โ โ s a a', equalPriorBayesError (K (s, a)) (K (s, a')) = 1/2.
โ {S : Type uS} {A : Type uA} {Y : Type uY} [inst : Fintype Y] (K : AuditCP.AuditChannel S A Y),
AuditCP.AuditSealed K โ โ (s : S) (a a' : A), AuditCP.equalPriorBayesError (K (s, a)) (K (s, a')) = 1 / 2โ {S : Type uS} {A : Type uA} {Y : Type uY} [inst : Fintype Y] (K : AuditCP.AuditChannel S A Y),
AuditCP.AuditSealed K โ โ (s : S) (a a' : A), AuditCP.equalPriorBayesError (K (s, a)) (K (s, a')) = 1 / 2F3b โ equalPriorBayesError p q = 1/2 โ p = q.
โ {Y : Type uY} [inst : Fintype Y] (p q : PMF Y), AuditCP.equalPriorBayesError p q = 1 / 2 โ p = qfiniteTotalVariation p q = 0 โ p = q.
โ {Y : Type uY} [inst : Fintype Y] (p q : PMF Y), AuditCP.finiteTotalVariation p q = 0 โ p = qF3a โ equalPriorBayesError p q = (1 - finiteTotalVariation p q) / 2.
โ {Y : Type uY} [inst : Fintype Y] (p q : PMF Y),
AuditCP.equalPriorBayesError p q = (1 - AuditCP.finiteTotalVariation p q) / 2On a finite type, the real point masses of p sum to 1.
โ {Y : Type uY} [inst : Fintype Y] (p : PMF Y), โ y, AuditCP.pmfMass p y = 1min x y = (x + y - |x - y|) / 2.
โ (x y : โ), min x y = (x + y - |x - y|) / 2
B6 โ AuditChannel S A Y is AuditExperiment (S ร A) Y, an experiment whose parameter is split into a public coordinate S and an audit coordinate A.
AuditCP.AuditIndex โ AuditCP.AuditIndex โ AuditCP.AuditIndex โ Type (max u v w)
The channel K is audit-sealed: its output law at fixed public state s does not depend on the audit state a.
{S : Type uS} โ {A : Type uA} โ {Y : Type uY} โ AuditCP.AuditChannel S A Y โ PropThe equal-prior Bayes error between p and q, (1/2) * โ y, min (pmfMass p y) (pmfMass q y).
{Y : Type uY} โ [Fintype Y] โ PMF Y โ PMF Y โ โThe total variation distance between p and q, half the โยน distance of their point masses, (1/2) * โ y, |pmfMass p y - pmfMass q y|.
{Y : Type uY} โ [Fintype Y] โ PMF Y โ PMF Y โ โThe real-valued point mass of y under the Mathlib PMF p, (p y).toReal.
{Y : Type uY} โ PMF Y โ Y โ โChoquet capacitability. For analytic sets, capacity equals the supremum over compact subsets (Kechris 30.13).
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
{cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ
โ {s : Set ฮฑ}, MeasureTheory.AnalyticSet s โ cap s = โจ K, โจ (_ : IsCompact K), โจ (_ : K โ s), cap Kโ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [inst_1 : MeasurableSpace ฮฑ] [BorelSpace ฮฑ] [PolishSpace ฮฑ]
{cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ
โ {s : Set ฮฑ}, MeasureTheory.AnalyticSet s โ cap s = โจ K, โจ (_ : IsCompact K), โจ (_ : K โ s), cap Kโ (N : โ โ โ) (n : โ), Monotone fun k => Cyl N n โฉ {g | g (n + 1) โค k}The intersection of closures of cylinder images equals the compact image. Key lemma for the capacitability proof: uses truncation and sequential compactness.
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] [PolishSpace ฮฑ] {f : (โ โ โ) โ ฮฑ},
Continuous f โ โ (N : โ โ โ), โ n, closure (f '' Cyl N n) = f '' Bnd Nโ (N g : โ โ โ), choquetTruncate N g โ Bnd N
โ (N : โ โ โ) (n : โ), โ g โ Cyl N n, โ i โค n, choquetTruncate N g i = g i
โ (N : โ โ โ), IsCompact (Bnd N)
โ (N : โ โ โ) (n : โ), Bnd N โ Cyl N n
โ (N : โ โ โ) (n : โ), Cyl N n = โ k, Cyl N n โฉ {g | g (n + 1) โค k}โ (N : โ โ โ) (n k : โ), Cyl N n โฉ {g | g (n + 1) โค k} = Cyl (Function.update N (n + 1) k) (n + 1)โ (N N' : โ โ โ) (n : โ), (โ i โค n, N i = N' i) โ Cyl N n = Cyl N' n
โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] {cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ โ {s t : Set ฮฑ}, s โ t โ cap s โค cap tโ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] {cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ โ (f : โ โ Set ฮฑ), Monotone f โ cap (โ n, f n) = โจ n, cap (f n)โ {ฮฑ : Type u_1} [inst : TopologicalSpace ฮฑ] {cap : Set ฮฑ โ ENNReal},
MeasureTheory.IsChoquetCapacity cap โ
โ (f : โ โ Set ฮฑ), Antitone f โ (โ (n : โ), IsClosed (f n)) โ cap (โ n, f n) = โจ
n, cap (f n)MeasureTheory.IsChoquetCapacity cap
MeasureTheory.AnalyticSet s
Bounded functions set: {g : โ โ โ | โ i, g i โค N i}.
(โ โ โ) โ Set (โ โ โ)
Cylinder set: {g : โ โ โ | โ i โค n, g i โค N i}.
(โ โ โ) โ โ โ Set (โ โ โ)
Bundled record of the three Choquet capacity axioms: monotonicity, sequential continuity from below along increasing unions, and sequential continuity from above along decreasing intersections of closed sets. The third axiom distinguishes a capacity from a general outer measure.
{ฮฑ : Type u_1} โ [TopologicalSpace ฮฑ] โ (Set ฮฑ โ ENNReal) โ PropTruncation: replace g i by min (g i) (N i) to bring any g into the bounded set.
(โ โ โ) โ (โ โ โ) โ โ โ โ
The discrete Vorob'ev theorem: a cover's base structure is acyclic exactly when local pairwise consistency always forces a global structure, i.e. every pairwise-consistent family of nonempty local relations on it glues.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] (๐ : Finset (Finset ฮน)),
HiddenChannelCapacity.Acyclic ๐ โ
โ (fam : List (HiddenChannelCapacity.LocalStruct ฮน fun x => ZMod 2)),
List.map (fun x => x.team) fam = ๐.toList โ
HiddenChannelCapacity.Consistent fam โ (โ L โ fam, L.rel.Nonempty) โ HiddenChannelCapacity.Glues famโ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] (๐ : Finset (Finset ฮน)),
HiddenChannelCapacity.Acyclic ๐ โ
โ (fam : List (HiddenChannelCapacity.LocalStruct ฮน fun x => ZMod 2)),
List.map (fun x => x.team) fam = ๐.toList โ
HiddenChannelCapacity.Consistent fam โ (โ L โ fam, L.rel.Nonempty) โ HiddenChannelCapacity.Glues famThe order-free amalgamation theorem: an acyclic cover forces every pairwise-consistent family with duplicate-free teams to glue. The base structure, acyclicity, is decidable, so the sufficiency is machine-checkable.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {X : ฮน โ Type u_2} [inst_1 : (i : ฮน) โ DecidableEq (X i)] [inst_2 : Fintype ฮน]
[(i : ฮน) โ Fintype (X i)] [โ (i : ฮน), Nonempty (X i)] (fam : List (HiddenChannelCapacity.LocalStruct ฮน X)),
HiddenChannelCapacity.Consistent fam โ
(List.map (fun x => x.team) fam).Nodup โ
HiddenChannelCapacity.Acyclic (List.map (fun x => x.team) fam).toFinset โ HiddenChannelCapacity.Glues famThe family-level running intersection property reads only the teams.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {X : ฮน โ Type u_2} (fam : List (HiddenChannelCapacity.LocalStruct ฮน X)),
HiddenChannelCapacity.RIP fam โ HiddenChannelCapacity.RIPCover (List.map (fun x => x.team) fam)โ {ฮน : Type u_1} {X : ฮน โ Type u_2} [inst : DecidableEq ฮน] (fam : List (HiddenChannelCapacity.LocalStruct ฮน X)),
HiddenChannelCapacity.famCover fam = HiddenChannelCapacity.listCover (List.map (fun x => x.team) fam)The amalgamation leg: over a running-intersection cover, a pairwise-consistent family glues: the base geometry forces the global structure.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {X : ฮน โ Type u_2} [inst_1 : (i : ฮน) โ DecidableEq (X i)] [inst_2 : Fintype ฮน]
[(i : ฮน) โ Fintype (X i)] [โ (i : ฮน), Nonempty (X i)] (fam : List (HiddenChannelCapacity.LocalStruct ฮน X)),
HiddenChannelCapacity.RIP fam โ HiddenChannelCapacity.Consistent fam โ HiddenChannelCapacity.Glues famThe amalgamation theorem (positive Vorob'ev, certificate form): once the cover carries the running intersection property, the natural join of any pairwise-consistent family realizes every member exactly, so the base geometry licenses the global structure.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {X : ฮน โ Type u_2} [inst_1 : (i : ฮน) โ DecidableEq (X i)] [inst_2 : Fintype ฮน]
[inst_3 : (i : ฮน) โ Fintype (X i)] [โ (i : ฮน), Nonempty (X i)] (fam : List (HiddenChannelCapacity.LocalStruct ฮน X)),
HiddenChannelCapacity.RIP fam โ
HiddenChannelCapacity.Consistent fam โ
โ L โ fam, HiddenChannelCapacity.projS (HiddenChannelCapacity.joinFam fam) L.team = L.relโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {X : ฮน โ Type u_2} {fam : List (HiddenChannelCapacity.LocalStruct ฮน X)}
{L : HiddenChannelCapacity.LocalStruct ฮน X}, L โ fam โ L.team โ HiddenChannelCapacity.famCover famThe presheaf law: the restriction map carries the T-marginal onto the S-marginal. This is the compatibility any ฤech-type complex over the marginal family consumes.
โ {ฮน : Type u_1} {X : ฮน โ Type u_2} [inst : (i : ฮน) โ DecidableEq (X i)] (ฮบ : Finset ((i : ฮน) โ X i)) {S T : Finset ฮน}
(hST : S โ T), Finset.image (Finset.restrictโ hST) (HiddenChannelCapacity.projS ฮบ T) = HiddenChannelCapacity.projS ฮบ Sโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {X : ฮน โ Type u_2} [inst_1 : (i : ฮน) โ DecidableEq (X i)] [inst_2 : Fintype ฮน]
[inst_3 : (i : ฮน) โ Fintype (X i)] {fam : List (HiddenChannelCapacity.LocalStruct ฮน X)} {f : (i : ฮน) โ X i},
f โ HiddenChannelCapacity.joinFam fam โ โ L โ fam, L.team.restrict f โ L.relAny local configuration extends to a global one, on nonempty parts.
โ {ฮน : Type u_1} {X : ฮน โ Type u_2} [โ (i : ฮน), Nonempty (X i)] (T : Finset ฮน) (g : (i : โฅT) โ X โi),
โ f, T.restrict f = gSub-marginalizing an interface marginal composes.
โ {ฮน : Type u_1} {X : ฮน โ Type u_2} [inst : (i : ฮน) โ DecidableEq (X i)] (L : HiddenChannelCapacity.LocalStruct ฮน X)
{S T : Finset ฮน} (hS : S โ T) (hT : T โ L.team), Finset.image (Finset.restrictโ hS) (L.imarg hT) = L.imarg โฏGluing is permutation-invariant.
โ {ฮน : Type u_1} {X : ฮน โ Type u_2} [inst : (i : ฮน) โ DecidableEq (X i)] [inst_1 : Fintype ฮน] [(i : ฮน) โ Fintype (X i)]
{fam fam' : List (HiddenChannelCapacity.LocalStruct ฮน X)},
fam'.Perm fam โ HiddenChannelCapacity.Glues fam' โ HiddenChannelCapacity.Glues famA permutation of a mapped list lifts along the map.
โ {ฮฑ : Type u_3} {ฮฒ : Type u_4} {g : ฮฑ โ ฮฒ} {ฯ : List ฮฒ} {l : List ฮฑ},
ฯ.Perm (List.map g l) โ โ l', l'.Perm l โง List.map g l' = ฯPairwise consistency is permutation-invariant.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {X : ฮน โ Type u_2} [inst_1 : (i : ฮน) โ DecidableEq (X i)]
{fam fam' : List (HiddenChannelCapacity.LocalStruct ฮน X)},
fam'.Perm fam โ HiddenChannelCapacity.Consistent fam โ HiddenChannelCapacity.Consistent fam'The CAPSTONE: a non-acyclic cover hosts a pairwise-consistent family of nonempty local relations that no global structure realizes: the base structure carries what local consistency cannot.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] (๐ : Finset (Finset ฮน)),
ยฌHiddenChannelCapacity.Acyclic ๐ โ
โ fam,
List.map (fun x => x.team) fam = ๐.toList โง
HiddenChannelCapacity.Consistent fam โง (โ L โ fam, L.rel.Nonempty) โง ยฌHiddenChannelCapacity.Glues famThe ear-free dichotomy: an ear-free cover is non-conformal or contains a chordless cycle. Dirac's lemma on the shared vertices produces a simplicial shared vertex, and the one-step GYO collapse turns it into an ear.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (๐ : Finset (Finset ฮน)),
HiddenChannelCapacity.EarFree ๐ โ ยฌHiddenChannelCapacity.Conformal ๐ โจ โ l, HiddenChannelCapacity.IsChordlessCycle ๐ lAn ear-free cover has a shared vertex: every team has nonempty boundary.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)},
HiddenChannelCapacity.EarFree ๐ โ (HiddenChannelCapacity.shared ๐).NonemptyThe one-step GYO collapse: in a conformal cover, a simplicial shared vertex yields an ear โ every team containing it has boundary inside the team covering its closed shared neighborhood.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)},
HiddenChannelCapacity.Conformal ๐ โ
โ {v : ฮน},
v โ HiddenChannelCapacity.shared ๐ โ
(โ x โ HiddenChannelCapacity.shared ๐,
โ y โ HiddenChannelCapacity.shared ๐,
HiddenChannelCapacity.Adj ๐ v x โ
HiddenChannelCapacity.Adj ๐ v y โ x โ y โ HiddenChannelCapacity.Adj ๐ x y) โ
โ T โ ๐, HiddenChannelCapacity.IsEarOf ๐ TA team's boundary is its shared part.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {T : Finset ฮน},
T โ ๐ โ T โฉ HiddenChannelCapacity.coverU (๐.erase T) = T โฉ HiddenChannelCapacity.shared ๐โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {v : ฮน},
v โ HiddenChannelCapacity.shared ๐ โ โ T โ ๐, โ T' โ ๐, T โ T' โง v โ T โง v โ T'Dirac's lemma, strong form: in a chordless-cycle-free adjacency structure, every vertex set is a clique or contains two nonadjacent simplicial vertices.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (๐ : Finset (Finset ฮน)),
(ยฌโ l, HiddenChannelCapacity.IsChordlessCycle ๐ l) โ
โ (S : Finset ฮน),
(โ x โ S, โ y โ S, x โ y โ HiddenChannelCapacity.Adj ๐ x y) โจ
โ u w,
HiddenChannelCapacity.SimpIn ๐ S u โง
HiddenChannelCapacity.SimpIn ๐ S w โง u โ w โง ยฌHiddenChannelCapacity.Adj ๐ u wThe assembled-cycle contradiction: two length-minimal x-y connectors through disjoint, mutually non-adjacent sides cannot coexist with chordless-cycle-freeness when x and y are non-adjacent.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)},
(ยฌโ l, HiddenChannelCapacity.IsChordlessCycle ๐ l) โ
โ {QA QB : ฮน โ Prop} {x y : ฮน},
x โ y โ
ยฌHiddenChannelCapacity.Adj ๐ x y โ
ยฌQB x โ
ยฌQB y โ
(โ (v : ฮน), QA v โ QB v โ False) โ
(โ (u v : ฮน), QA u โ QB v โ ยฌHiddenChannelCapacity.Adj ๐ u v) โ
โ (pA : List ฮน),
List.IsChain (HiddenChannelCapacity.Adj ๐) pA โ
pA.Nodup โ
pA.head? = some x โ
pA.getLast? = some y โ
(โ v โ pA, v = x โจ v = y โจ QA v) โ
(โ (q : List ฮน),
(List.IsChain (HiddenChannelCapacity.Adj ๐) q โง
q.Nodup โง
q.head? = some x โง q.getLast? = some y โง โ v โ q, v = x โจ v = y โจ QA v) โ
pA.length โค q.length) โ
โ (pB : List ฮน),
List.IsChain (HiddenChannelCapacity.Adj ๐) pB โ
pB.Nodup โ
pB.head? = some x โ
pB.getLast? = some y โ
(โ v โ pB, v = x โจ v = y โจ QB v) โ
(โ (q : List ฮน),
(List.IsChain (HiddenChannelCapacity.Adj ๐) q โง
q.Nodup โง
q.head? = some x โง
q.getLast? = some y โง โ v โ q, v = x โจ v = y โจ QB v) โ
pB.length โค q.length) โ
FalseMinimality kills chords: a length-minimal constrained x-y path has no adjacency between positions at distance โฅ 2.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {Q : ฮน โ Prop} {x y : ฮน} {p : List ฮน},
List.IsChain (HiddenChannelCapacity.Adj ๐) p โ
p.head? = some x โ
p.getLast? = some y โ
(โ v โ p, v = x โจ v = y โจ Q v) โ
(โ (q : List ฮน),
(List.IsChain (HiddenChannelCapacity.Adj ๐) q โง
q.Nodup โง q.head? = some x โง q.getLast? = some y โง โ v โ q, v = x โจ v = y โจ Q v) โ
p.length โค q.length) โ
p.Nodup โ
โ {i j : โ} (hi : i < p.length) (hj : j < p.length), i + 2 โค j โ ยฌHiddenChannelCapacity.Adj ๐ p[i] p[j]โ {ฮน : Type u_1} {p : List ฮน} {y : ฮน}, p.getLast? = some y โ โ (h0 : 0 < p.length), p[p.length - 1] = yโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v w : ฮน},
HiddenChannelCapacity.ReachOn ๐ P u v โ HiddenChannelCapacity.ReachOn ๐ P v w โ HiddenChannelCapacity.ReachOn ๐ P u wโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v : ฮน},
HiddenChannelCapacity.ReachOn ๐ P u v โ HiddenChannelCapacity.ReachOn ๐ P v uโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u : ฮน}, P u โ HiddenChannelCapacity.ReachOn ๐ P u uEvery vertex on a walk is reachable from its head.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v : ฮน} {p : List ฮน},
HiddenChannelCapacity.IsWalkOn ๐ P p โ p.head? = some u โ v โ p โ HiddenChannelCapacity.ReachOn ๐ P u vโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v w : ฮน},
HiddenChannelCapacity.ReachOn ๐ P u v โ HiddenChannelCapacity.Adj ๐ v w โ P w โ HiddenChannelCapacity.ReachOn ๐ P u wโ {ฮน : Type u_1} {Q : List ฮน โ Prop}, (โ p, Q p) โ โ p, Q p โง โ (q : List ฮน), Q q โ p.length โค q.lengthA connector: an x-y path with interior inside Q, from adjacent entry points and interior reachability.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {Q : ฮน โ Prop} {x y uโ uโ : ฮน},
x โ y โ
HiddenChannelCapacity.Adj ๐ x uโ โ
HiddenChannelCapacity.Adj ๐ y uโ โ
HiddenChannelCapacity.ReachOn ๐ Q uโ uโ โ
ยฌQ x โ
ยฌQ y โ
โ p,
List.IsChain (HiddenChannelCapacity.Adj ๐) p โง
p.Nodup โง p.head? = some x โง p.getLast? = some y โง โ v โ p, v = x โจ v = y โจ Q vIf x has no neighbor reachable from a off the separator, then erasing x still separates: the first x-occurrence on any violating path has its predecessor reachable from a and adjacent to x.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {S Sep : Finset ฮน} {a b x : ฮน},
x โ a โ
(โ (u : ฮน),
u โ S \ Sep โง HiddenChannelCapacity.ReachOn ๐ (fun v => v โ S \ Sep) a u โ ยฌHiddenChannelCapacity.Adj ๐ x u) โ
ยฌHiddenChannelCapacity.ReachOn ๐ (fun v => v โ S \ Sep) a b โ
ยฌHiddenChannelCapacity.ReachOn ๐ (fun v => v โ S \ Sep.erase x) a bEvery chained walk contains a duplicate-free path with the same endpoints and vertices.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} (p : List ฮน),
HiddenChannelCapacity.IsWalkOn ๐ P p โ
โ {u v : ฮน}, p.head? = some u โ p.getLast? = some v โ HiddenChannelCapacity.ReachOn ๐ P u vโ {ฮน : Type u_1} {R : ฮน โ ฮน โ Prop} {p : List ฮน}, List.IsChain R p โ โ (n : โ), List.IsChain R (List.drop n p)โ {ฮน : Type u_1} {R : ฮน โ ฮน โ Prop} {p : List ฮน}, List.IsChain R p โ โ (n : โ), List.IsChain R (List.take n p)Positional form of the head.
โ {ฮน : Type u_1} {p : List ฮน} {x : ฮน}, p.head? = some x โ โ (h0 : 0 < p.length), p[0] = xโ {ฮน : Type u_1} (l : List ฮน) {s t : โ} (hs : s < l.length) (ht : t < l.length), s = t โ l[s] = l[t]โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {u v : ฮน}, HiddenChannelCapacity.Adj ๐ u v โ HiddenChannelCapacity.Adj ๐ v uThe capstone, modulo the dichotomy: once every ear-free core is non-conformal or has a chordless cycle (the Dirac edge, URS D26 L6), every non-Graham-reducible cover hosts the base structure that licenses the verdict: a pairwise-consistent family of nonempty local relations that no global structure realizes.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] (๐ : Finset (Finset ฮน)),
ยฌHiddenChannelCapacity.GrahamReducible ๐ โ
(โ ๐' โ ๐,
HiddenChannelCapacity.EarFree ๐' โ
ยฌHiddenChannelCapacity.Conformal ๐' โจ โ l, HiddenChannelCapacity.IsChordlessCycle ๐' l) โ
โ fam,
List.map (fun x => x.team) fam = ๐.toList โง
HiddenChannelCapacity.Consistent fam โง (โ L โ fam, L.rel.Nonempty) โง ยฌHiddenChannelCapacity.Glues famChain extraction: a non-Graham-reducible cover has an ear-free core such that any vertex set inside the core's cover hosted anywhere in the cover is hosted inside the core.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (๐ : Finset (Finset ฮน)),
ยฌHiddenChannelCapacity.GrahamReducible ๐ โ
โ ๐' โ ๐,
HiddenChannelCapacity.EarFree ๐' โง โ w โ HiddenChannelCapacity.coverU ๐', (โ V โ ๐, w โ V) โ โ W โ ๐', w โ WThe clique datum: a cardinality-minimal non-covered clique yields a clean system. Minimality covers the erase-sets; non-coveredness is cleanliness; parity closes by counting.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (๐ : Finset (Finset ฮน)),
๐.Nonempty โ
ยฌHiddenChannelCapacity.Conformal ๐ โ
โ ๐ฒ,
๐ฒ.Nonempty โง
(โ w โ ๐ฒ, w โ โ
) โง
(โ w โ ๐ฒ, โ T โ ๐, w โ T) โง
(โ V โ ๐, โ wโ โ ๐ฒ, โ wโ โ ๐ฒ, wโ โ V โ wโ โ V โ wโ = wโ) โง HiddenChannelCapacity.xorSum ๐ฒ = โ
An F2 sum vanishes iff every incidence count is even.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {D : Finset (Finset ฮน)},
HiddenChannelCapacity.xorSum D = โ
โ โ (i : ฮน), ยฌOdd {x โ D | i โ x}.cardMembership in an F2 sum is odd incidence.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {D : Finset (Finset ฮน)} {i : ฮน},
i โ HiddenChannelCapacity.xorSum D โ Odd {x โ D | i โ x}.cardโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {K : Finset ฮน} {i j : ฮน}, i โ j โ K.erase i โช K.erase j = KSmall cliques are covered: the empty set (nonempty cover), singletons, edges.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)},
๐.Nonempty โ
โ {K : Finset ฮน},
K โ HiddenChannelCapacity.coverU ๐ โ
HiddenChannelCapacity.IsClique ๐ K โ K.card โค 2 โ HiddenChannelCapacity.Covered ๐ Kโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {v : ฮน},
v โ HiddenChannelCapacity.coverU ๐ โ โ T โ ๐, v โ TThe clean system of a chordless cycle, in finder form.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {l : List ฮน},
HiddenChannelCapacity.IsChordlessCycle ๐ l โ
โ ๐ฒ,
๐ฒ.Nonempty โง
(โ w โ ๐ฒ, w โ โ
) โง
(โ w โ ๐ฒ, โ T โ ๐, w โ T) โง
(โ V โ ๐, โ wโ โ ๐ฒ, โ wโ โ ๐ฒ, wโ โ V โ wโ โ V โ wโ = wโ) โง HiddenChannelCapacity.xorSum ๐ฒ = โ
The cycle sum telescopes: consecutive pairs of a duplicate-free cyclic list sum to โ
over F2 โ every vertex lies in exactly two pairs.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (l : List ฮน),
l.Nodup โ 2 โค l.length โ HiddenChannelCapacity.xorSumL (HiddenChannelCapacity.cyclePairs l) = โ
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (lโ lโ : List (Finset ฮน)),
lโ.length = lโ.length โ
HiddenChannelCapacity.xorSumL (List.zipWith (fun x1 x2 => symmDiff x1 x2) lโ lโ) =
symmDiff (HiddenChannelCapacity.xorSumL lโ) (HiddenChannelCapacity.xorSumL lโ)โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {lโ lโ : List (Finset ฮน)},
lโ.Perm lโ โ HiddenChannelCapacity.xorSumL lโ = HiddenChannelCapacity.xorSumL lโโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {a b : ฮน}, a โ b โ {a, b} = symmDiff {a} {b}โ {ฮน : Type u_1} (l : List ฮน), l.Nodup โ โ (h2 : 2 โค l.length) (i : โ) (hi : i < l.length), l[i] โ l[(i + 1) % l.length]Distinct positions carry distinct pairs: the pair list is duplicate-free.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (l : List ฮน),
l.Nodup โ 3 โค l.length โ (HiddenChannelCapacity.cyclePairs l).NodupMembers of the cycle-pair list are nonempty.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {l : List ฮน},
0 < l.length โ โ {w : Finset ฮน}, w โ HiddenChannelCapacity.cyclePairs l โ w โ โ
Chordlessness is cleanliness: no team hosts two distinct pairs of a chordless cycle.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {l : List ฮน},
HiddenChannelCapacity.IsChordlessCycle ๐ l โ
โ V โ ๐,
โ wโ โ HiddenChannelCapacity.cyclePairs l, โ wโ โ HiddenChannelCapacity.cyclePairs l, wโ โ V โ wโ โ V โ wโ = wโโ (n i : โ), i < n โ (i + 1) % n = i + 1 โง i + 1 < n โจ (i + 1) % n = 0 โง i + 1 = n
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (l : List ฮน) (h0 : 0 < l.length) (i : โ)
(h : i < (HiddenChannelCapacity.cyclePairs l).length),
(HiddenChannelCapacity.cyclePairs l)[i] = {l[i], l[(i + 1) % l.length]}โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (l : List ฮน), (HiddenChannelCapacity.cyclePairs l).length = l.lengthCo-hosted vertices are adjacent or equal.
โ {ฮน : Type u_1} [DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {V : Finset ฮน},
V โ ๐ โ โ {u v : ฮน}, u โ V โ v โ V โ HiddenChannelCapacity.Adj ๐ u v โจ u = vThe transfer: a clean system on a chained ear-free core is a clean system on the full cover โ any team hosting two members routes them into a single core team.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] (๐ ๐' : Finset (Finset ฮน)),
๐' โ ๐ โ
(โ w โ HiddenChannelCapacity.coverU ๐', (โ V โ ๐, w โ V) โ โ W โ ๐', w โ W) โ
โ (๐ฒ : Finset (Finset ฮน)),
๐ฒ.Nonempty โ
(โ w โ ๐ฒ, w โ โ
) โ
(โ w โ ๐ฒ, โ T โ ๐', w โ T) โ
(โ V โ ๐', โ wโ โ ๐ฒ, โ wโ โ ๐ฒ, wโ โ V โ wโ โ V โ wโ = wโ) โ
HiddenChannelCapacity.xorSum ๐ฒ = โ
โ
โ fam,
List.map (fun x => x.team) fam = ๐.toList โง
HiddenChannelCapacity.Consistent fam โง
(โ L โ fam, L.rel.Nonempty) โง ยฌHiddenChannelCapacity.Glues famThe clean-system obstruction, the unified consumption theorem of C-ENGINE: a base-side clean system (a nonempty hosted family of distinct nonempty supports with vanishing F2 sum, no team hosting two) licenses a verdict on local measurement, a pairwise-consistent family of nonempty local relations over the whole cover that no global structure realizes. Every witness rung (chordless cycles, even boundary data, landings, the H_k clique parities) differs only as a FINDER of such a base system.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] (๐ ๐ฒ : Finset (Finset ฮน)),
๐ฒ.Nonempty โ
(โ w โ ๐ฒ, w โ โ
) โ
(โ w โ ๐ฒ, โ V โ ๐, w โ V) โ
(โ V โ ๐, โ wโ โ ๐ฒ, โ wโ โ ๐ฒ, wโ โ V โ wโ โ V โ wโ = wโ) โ
HiddenChannelCapacity.xorSum ๐ฒ = โ
โ
โ fam,
List.map (fun x => x.team) fam = ๐.toList โง
HiddenChannelCapacity.Consistent fam โง (โ L โ fam, L.rel.Nonempty) โง ยฌHiddenChannelCapacity.Glues famThe affine obstruction engine, packaged: from a base-side hosted odd dependency (over a closed coherent system) follows a verdict on local measurement, a family of nonempty local relations that is pairwise consistent yet realizes no global structure.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] {teams : List (Finset ฮน)} {๐ฒ : Finset (Finset ฮน)}
{c : Finset ฮน โ ZMod 2},
teams โ [] โ
(โ V โ teams, HiddenChannelCapacity.ClosedAt ๐ฒ V) โ
(โ V โ teams, HiddenChannelCapacity.CoherentAt ๐ฒ c V) โ
โ D โ ๐ฒ,
HiddenChannelCapacity.xorSum D = โ
โ
โ w โ D, c w = 1 โ
(โ w โ D, โ V โ teams, w โ V) โ
HiddenChannelCapacity.Consistent (HiddenChannelCapacity.affFam teams ๐ฒ c) โง
(โ L โ HiddenChannelCapacity.affFam teams ๐ฒ c, L.rel.Nonempty) โง
ยฌHiddenChannelCapacity.Glues (HiddenChannelCapacity.affFam teams ๐ฒ c)The obstruction theorem: a hosted odd dependency (a base-side witness) forecloses every global structure; any global realizing all the local relations would satisfy every hosted constraint, and the dependency sums those satisfactions to 0 = 1.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] [inst_1 : Fintype ฮน] {teams : List (Finset ฮน)} {๐ฒ : Finset (Finset ฮน)}
{c : Finset ฮน โ ZMod 2},
teams โ [] โ
(โ V โ teams, HiddenChannelCapacity.CoherentAt ๐ฒ c V) โ
โ D โ ๐ฒ,
HiddenChannelCapacity.xorSum D = โ
โ
โ w โ D, c w = 1 โ
(โ w โ D, โ V โ teams, w โ V) โ ยฌHiddenChannelCapacity.Glues (HiddenChannelCapacity.affFam teams ๐ฒ c)โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (D : Finset (Finset ฮน)) (f : ฮน โ ZMod 2),
HiddenChannelCapacity.parityOn (HiddenChannelCapacity.xorSum D) f = โ w โ D, HiddenChannelCapacity.parityOn w fClosed coherent affine families are pairwise consistent: both interface marginals ARE the interface's visible system.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {teams : List (Finset ฮน)} {๐ฒ : Finset (Finset ฮน)} {c : Finset ฮน โ ZMod 2},
(โ V โ teams, HiddenChannelCapacity.ClosedAt ๐ฒ V) โ
(โ V โ teams, HiddenChannelCapacity.CoherentAt ๐ฒ c V) โ
HiddenChannelCapacity.Consistent (HiddenChannelCapacity.affFam teams ๐ฒ c)The exact marginal computation: the interface marginal of a closed coherent local structure is exactly the local structure of the interface โ the interface sees precisely the visible constraints.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V S : Finset ฮน} {๐ฒ : Finset (Finset ฮน)} {c : Finset ฮน โ ZMod 2} (hS : S โ V),
HiddenChannelCapacity.ClosedAt ๐ฒ V โ
HiddenChannelCapacity.CoherentAt ๐ฒ c V โ
(HiddenChannelCapacity.affLocal V ๐ฒ c).imarg hS = (HiddenChannelCapacity.affLocal S ๐ฒ c).relTransport of team parities along nested restriction: the value only reads the support.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V S w : Finset ฮน} (hS : S โ V),
w โ S โ โ (F : โฅV โ ZMod 2), HiddenChannelCapacity.wsum S w (Finset.restrictโ hS F) = HiddenChannelCapacity.wsum V w FThe pinned system: hosted constraints of V plus a singleton pin at every S-coordinate of a visible-constraint-satisfying trace. Coherent, by closedness.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V S : Finset ฮน} {๐ฒ : Finset (Finset ฮน)} {c : Finset ฮน โ ZMod 2},
HiddenChannelCapacity.ClosedAt ๐ฒ V โ
HiddenChannelCapacity.CoherentAt ๐ฒ c V โ
โ (gโ : ฮน โ ZMod 2),
(โ w โ ๐ฒ, w โ S โ HiddenChannelCapacity.parityOn w gโ = c w) โ
HiddenChannelCapacity.SysCoherent
(HiddenChannelCapacity.hostedListโ V ๐ฒ c ++ List.map (fun i => ({i}, gโ i)) S.toList)โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V : Finset ฮน} {๐ฒ D : Finset (Finset ฮน)},
D โ {x โ ๐ฒ | x โ V} โ HiddenChannelCapacity.xorSum D โ Vโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {l : List ฮน},
l.Nodup โ HiddenChannelCapacity.xorSumL (List.map (fun x => {x}) l) = l.toFinsetโ {ฮน : Type u_1} [inst : DecidableEq ฮน] (a b : List (Finset ฮน)),
HiddenChannelCapacity.xorSumL (a ++ b) = symmDiff (HiddenChannelCapacity.xorSumL a) (HiddenChannelCapacity.xorSumL b)โ {ฮน : Type u_1} [inst : DecidableEq ฮน], HiddenChannelCapacity.xorSumL [] = โ
โ {ฮน : Type u_1} (gโ : ฮน โ ZMod 2) (l : List ฮน), List.map Prod.snd (List.map (fun i => ({i}, gโ i)) l) = List.map gโ lโ {ฮน : Type u_1} (gโ : ฮน โ ZMod 2) (l : List ฮน),
List.map Prod.fst (List.map (fun i => ({i}, gโ i)) l) = List.map (fun x => {x}) lEvery local relation of a coherent system is nonempty.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V : Finset ฮน} {๐ฒ : Finset (Finset ฮน)} {c : Finset ฮน โ ZMod 2},
HiddenChannelCapacity.CoherentAt ๐ฒ c V โ (HiddenChannelCapacity.affLocal V ๐ฒ c).rel.Nonemptyโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V w : Finset ฮน},
w โ V โ โ (f : ฮน โ ZMod 2), HiddenChannelCapacity.wsum V w (V.restrict f) = HiddenChannelCapacity.parityOn w fโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V : Finset ฮน} {๐ฒ : Finset (Finset ฮน)} {c : Finset ฮน โ ZMod 2},
HiddenChannelCapacity.CoherentAt ๐ฒ c V โ HiddenChannelCapacity.SysCoherent (HiddenChannelCapacity.hostedListโ V ๐ฒ c)โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {l : List (Finset ฮน)},
l.Nodup โ HiddenChannelCapacity.xorSumL l = HiddenChannelCapacity.xorSum l.toFinsetโ {ฮน : Type u_1} (c : Finset ฮน โ ZMod 2) (l : List (Finset ฮน)),
List.map Prod.snd (List.map (fun w => (w, c w)) l) = List.map c lโ {ฮน : Type u_1} (c : Finset ฮน โ ZMod 2) (l : List (Finset ฮน)), List.map Prod.fst (List.map (fun w => (w, c w)) l) = lโ {ฮน : Type u_1} [inst : DecidableEq ฮน] {V : Finset ฮน} {๐ฒ : Finset (Finset ฮน)} {c : Finset ฮน โ ZMod 2}
{F : โฅV โ ZMod 2},
F โ (HiddenChannelCapacity.affLocal V ๐ฒ c).rel โ โ w โ ๐ฒ, w โ V โ HiddenChannelCapacity.wsum V w F = c wThe solvability core: a coherent list-presented parity system has a solution. Elementary Gaussian elimination, by induction on the list.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (L : List (Finset ฮน ร ZMod 2)),
HiddenChannelCapacity.SysCoherent L โ โ f, โ p โ L, HiddenChannelCapacity.parityOn p.1 f = p.2โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (w : Finset ฮน) (l : List (Finset ฮน)),
HiddenChannelCapacity.xorSumL (w :: l) = symmDiff w (HiddenChannelCapacity.xorSumL l)The Gaussian transform bookkeeping: transforming the x-containing entries by โ wโ shifts the F2 sum by wโ per transformed entry, and the constants by cโ likewise.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (x : ฮน) (wโ : Finset ฮน) (cโ : ZMod 2) (sub : List (Finset ฮน ร ZMod 2)),
HiddenChannelCapacity.xorSumL
(List.map Prod.fst (List.map (fun p => if x โ p.1 then (symmDiff p.1 wโ, p.2 + cโ) else p) sub)) =
symmDiff (HiddenChannelCapacity.xorSumL (List.map Prod.fst sub))
(if List.countP (fun p => decide (x โ p.1)) sub % 2 = 1 then wโ else โ
) โง
(List.map Prod.snd (List.map (fun p => if x โ p.1 then (symmDiff p.1 wโ, p.2 + cโ) else p) sub)).sum =
(List.map Prod.snd sub).sum + if List.countP (fun p => decide (x โ p.1)) sub % 2 = 1 then cโ else 0โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (a : Finset ฮน), symmDiff a โ
= aโ {ฮน : Type u_1} [inst : DecidableEq ฮน] (w w' : Finset ฮน) (f : ฮน โ ZMod 2),
HiddenChannelCapacity.parityOn (symmDiff w w') f =
HiddenChannelCapacity.parityOn w f + HiddenChannelCapacity.parityOn w' fโ (a : ZMod 2), a + a = 0
The Graham bridge: acyclicity is exactly Graham reducibility.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (๐ : Finset (Finset ฮน)),
HiddenChannelCapacity.Acyclic ๐ โ HiddenChannelCapacity.GrahamReducible ๐An RIP order eliminates from the front: its head is an ear and the tail remains RIP.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (order : List (Finset ฮน)),
order.Nodup โ HiddenChannelCapacity.RIPCover order โ HiddenChannelCapacity.GrahamN order.length order.toFinsetAn elimination run is a running-intersection order read forward.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (n : โ) (๐ : Finset (Finset ฮน)),
HiddenChannelCapacity.GrahamN n ๐ โ โ order, order.Nodup โง order.toFinset = ๐ โง HiddenChannelCapacity.RIPCover orderโ {ฮน : Type u_1} [inst : DecidableEq ฮน] (l : List (Finset ฮน)),
HiddenChannelCapacity.listCover l = HiddenChannelCapacity.coverU l.toFinsetOrder-free acyclicity: some duplicate-free enumeration of the cover is a running-intersection order.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ PropPrimal adjacency: distinct co-hosted vertices.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ ฮน โ ฮน โ PropIntra-team span closure: subset sums of the constraints hosted by V vanish or stay in the system.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ Finset ฮน โ PropIntra-team coherence: constants vanish on the dependencies hosted by V.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ (Finset ฮน โ ZMod 2) โ Finset ฮน โ PropConformality: every clique inside the cover is covered.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ PropPairwise consistency: any two members expose the same interface marginal.
{ฮน : Type u_1} โ
[DecidableEq ฮน] โ
{X : ฮน โ Type u_2} โ [(i : ฮน) โ DecidableEq (X i)] โ List (HiddenChannelCapacity.LocalStruct ฮน X) โ PropA vertex set covered by a single team.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ Finset ฮน โ PropAn ear-free nonempty cover: the stuck core condition.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ PropGluing: some global structure realizes every member exactly.
{ฮน : Type u_1} โ
{X : ฮน โ Type u_2} โ [(i : ฮน) โ DecidableEq (X i)] โ [Fintype ฮน] โ List (HiddenChannelCapacity.LocalStruct ฮน X) โ PropNondeterministic ear elimination, fueled.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ โ โ Finset (Finset ฮน) โ PropGraham reducibility: ear elimination empties the cover; the cover's size is fuel enough since every elimination removes exactly one team.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ PropA chordless cycle: duplicate-free, length โฅ 4, consecutive pairs hosted, and the only adjacencies among its vertices are the cyclically consecutive ones.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ List ฮน โ PropA clique of the primal graph.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ Finset ฮน โ PropThe ear condition: the team meets the union of the others inside a single other team (or there are no others).
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ Finset ฮน โ PropA walk constrained to P: chained adjacencies, every vertex satisfies P.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ (ฮน โ Prop) โ List ฮน โ PropA local structure: a team of parts together with an ignorance structure on the team's marginal type.
(ฮน : Type u_3) โ (ฮน โ Type u_4) โ Type (max u_3 u_4)
The interface marginal of a local structure over a sub-team.
{ฮน : Type u_1} โ
{X : ฮน โ Type u_2} โ
[(i : ฮน) โ DecidableEq (X i)] โ
(L : HiddenChannelCapacity.LocalStruct ฮน X) โ {S : Finset ฮน} โ S โ L.team โ Finset ((i : โฅS) โ X โi)The admissible local configurations.
{ฮน : Type u_3} โ {X : ฮน โ Type u_4} โ (self : HiddenChannelCapacity.LocalStruct ฮน X) โ Finset ((i : โฅself.team) โ X โi)The team of parts this local structure constrains.
{ฮน : Type u_3} โ {X : ฮน โ Type u_4} โ HiddenChannelCapacity.LocalStruct ฮน X โ Finset ฮนRunning intersection property: each member meets the union of the later members inside a single later member (the head is the last-eliminated team).
{ฮน : Type u_1} โ [DecidableEq ฮน] โ {X : ฮน โ Type u_2} โ List (HiddenChannelCapacity.LocalStruct ฮน X) โ PropTeam-level running intersection property: each team meets the union of the later teams inside a single later team.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ List (Finset ฮน) โ PropReachability within P: a duplicate-free walk with the given endpoints.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ (ฮน โ Prop) โ ฮน โ ฮน โ Propv is simplicial within the vertex set S: it lies in S and its ๐-neighbors inside S are pairwise adjacent.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ Finset ฮน โ ฮน โ PropCoherence of a list-presented parity system: constants vanish on every dependency.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ List (Finset ฮน ร ZMod 2) โ PropThe affine family over a cover: every team with its fitting constraints.
{ฮน : Type u_1} โ
[DecidableEq ฮน] โ
List (Finset ฮน) โ
Finset (Finset ฮน) โ (Finset ฮน โ ZMod 2) โ List (HiddenChannelCapacity.LocalStruct ฮน fun x => ZMod 2)The team V with every constraint of the system that fits inside it: the uniform replication rule.
{ฮน : Type u_1} โ
[DecidableEq ฮน] โ
Finset ฮน โ Finset (Finset ฮน) โ (Finset ฮน โ ZMod 2) โ HiddenChannelCapacity.LocalStruct ฮน fun x => ZMod 2The union of a finite set of teams.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ Finset ฮนThe consecutive-pair supports of a cyclic vertex list.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ List ฮน โ List (Finset ฮน)Acyclicity is decidable โ the payoff of the bridge.
{ฮน : Type u_1} โ [inst : DecidableEq ฮน] โ (๐ : Finset (Finset ฮน)) โ Decidable (HiddenChannelCapacity.Acyclic ๐){ฮน : Type u_1} โ
[inst : DecidableEq ฮน] โ (๐ : Finset (Finset ฮน)) โ (T : Finset ฮน) โ Decidable (HiddenChannelCapacity.IsEarOf ๐ T)The union of a family's teams.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ {X : ฮน โ Type u_2} โ List (HiddenChannelCapacity.LocalStruct ฮน X) โ Finset ฮนThe hosted constraints of V, as a list system.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset ฮน โ Finset (Finset ฮน) โ (Finset ฮน โ ZMod 2) โ List (Finset ฮน ร ZMod 2){ฮน : Type u_1} โ [DecidableEq ฮน] โ (๐ : Finset (Finset ฮน)) โ (u v : ฮน) โ Decidable (HiddenChannelCapacity.Adj ๐ u v){ฮน : Type u_1} โ
[DecidableEq ฮน] โ (๐ : Finset (Finset ฮน)) โ (K : Finset ฮน) โ Decidable (HiddenChannelCapacity.Covered ๐ K){ฮน : Type u_1} โ
[DecidableEq ฮน] โ (๐ : Finset (Finset ฮน)) โ (K : Finset ฮน) โ Decidable (HiddenChannelCapacity.IsClique ๐ K)The natural join of a family: all global configurations passing every local test.
{ฮน : Type u_1} โ
[DecidableEq ฮน] โ
{X : ฮน โ Type u_2} โ
[(i : ฮน) โ DecidableEq (X i)] โ
[Fintype ฮน] โ [(i : ฮน) โ Fintype (X i)] โ List (HiddenChannelCapacity.LocalStruct ฮน X) โ Finset ((i : ฮน) โ X i)The union of a list of teams.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ List (Finset ฮน) โ Finset ฮนParity of a global configuration over a support.
{ฮน : Type u_1} โ Finset ฮน โ (ฮน โ ZMod 2) โ ZMod 2The marginal of a multipartite relation over a sub-team S: the image under Finset.restrict.
{ฮน : Type u_1} โ
{X : ฮน โ Type u_2} โ
[(i : ฮน) โ DecidableEq (X i)] โ Finset ((i : ฮน) โ X i) โ (S : Finset ฮน) โ Finset ((i : โฅS) โ X โi)The shared vertices of a cover.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ Finset ฮนParity of a team configuration over the team part of a support.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ (V : Finset ฮน) โ Finset ฮน โ (โฅV โ ZMod 2) โ ZMod 2F2 sum of a finite collection of supports.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ Finset ฮนF2 sum of a list of supports.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ List (Finset ฮน) โ Finset ฮนDirac's lemma, strong form: in a chordless-cycle-free adjacency structure, every vertex set is a clique or contains two nonadjacent simplicial vertices.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (๐ : Finset (Finset ฮน)),
(ยฌโ l, HiddenChannelCapacity.IsChordlessCycle ๐ l) โ
โ (S : Finset ฮน),
(โ x โ S, โ y โ S, x โ y โ HiddenChannelCapacity.Adj ๐ x y) โจ
โ u w,
HiddenChannelCapacity.SimpIn ๐ S u โง
HiddenChannelCapacity.SimpIn ๐ S w โง u โ w โง ยฌHiddenChannelCapacity.Adj ๐ u wโ {ฮน : Type u_1} [inst : DecidableEq ฮน] (๐ : Finset (Finset ฮน)),
(ยฌโ l, HiddenChannelCapacity.IsChordlessCycle ๐ l) โ
โ (S : Finset ฮน),
(โ x โ S, โ y โ S, x โ y โ HiddenChannelCapacity.Adj ๐ x y) โจ
โ u w,
HiddenChannelCapacity.SimpIn ๐ S u โง
HiddenChannelCapacity.SimpIn ๐ S w โง u โ w โง ยฌHiddenChannelCapacity.Adj ๐ u wThe assembled-cycle contradiction: two length-minimal x-y connectors through disjoint, mutually non-adjacent sides cannot coexist with chordless-cycle-freeness when x and y are non-adjacent.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)},
(ยฌโ l, HiddenChannelCapacity.IsChordlessCycle ๐ l) โ
โ {QA QB : ฮน โ Prop} {x y : ฮน},
x โ y โ
ยฌHiddenChannelCapacity.Adj ๐ x y โ
ยฌQB x โ
ยฌQB y โ
(โ (v : ฮน), QA v โ QB v โ False) โ
(โ (u v : ฮน), QA u โ QB v โ ยฌHiddenChannelCapacity.Adj ๐ u v) โ
โ (pA : List ฮน),
List.IsChain (HiddenChannelCapacity.Adj ๐) pA โ
pA.Nodup โ
pA.head? = some x โ
pA.getLast? = some y โ
(โ v โ pA, v = x โจ v = y โจ QA v) โ
(โ (q : List ฮน),
(List.IsChain (HiddenChannelCapacity.Adj ๐) q โง
q.Nodup โง
q.head? = some x โง q.getLast? = some y โง โ v โ q, v = x โจ v = y โจ QA v) โ
pA.length โค q.length) โ
โ (pB : List ฮน),
List.IsChain (HiddenChannelCapacity.Adj ๐) pB โ
pB.Nodup โ
pB.head? = some x โ
pB.getLast? = some y โ
(โ v โ pB, v = x โจ v = y โจ QB v) โ
(โ (q : List ฮน),
(List.IsChain (HiddenChannelCapacity.Adj ๐) q โง
q.Nodup โง
q.head? = some x โง
q.getLast? = some y โง โ v โ q, v = x โจ v = y โจ QB v) โ
pB.length โค q.length) โ
FalseMinimality kills chords: a length-minimal constrained x-y path has no adjacency between positions at distance โฅ 2.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {Q : ฮน โ Prop} {x y : ฮน} {p : List ฮน},
List.IsChain (HiddenChannelCapacity.Adj ๐) p โ
p.head? = some x โ
p.getLast? = some y โ
(โ v โ p, v = x โจ v = y โจ Q v) โ
(โ (q : List ฮน),
(List.IsChain (HiddenChannelCapacity.Adj ๐) q โง
q.Nodup โง q.head? = some x โง q.getLast? = some y โง โ v โ q, v = x โจ v = y โจ Q v) โ
p.length โค q.length) โ
p.Nodup โ
โ {i j : โ} (hi : i < p.length) (hj : j < p.length), i + 2 โค j โ ยฌHiddenChannelCapacity.Adj ๐ p[i] p[j]โ (n i : โ), i < n โ (i + 1) % n = i + 1 โง i + 1 < n โจ (i + 1) % n = 0 โง i + 1 = n
โ {ฮน : Type u_1} {p : List ฮน} {y : ฮน}, p.getLast? = some y โ โ (h0 : 0 < p.length), p[p.length - 1] = yโ {ฮน : Type u_1} [inst : DecidableEq ฮน] (l : List ฮน) (h0 : 0 < l.length) (i : โ)
(h : i < (HiddenChannelCapacity.cyclePairs l).length),
(HiddenChannelCapacity.cyclePairs l)[i] = {l[i], l[(i + 1) % l.length]}โ {ฮน : Type u_1} [inst : DecidableEq ฮน] (l : List ฮน), (HiddenChannelCapacity.cyclePairs l).length = l.lengthโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v w : ฮน},
HiddenChannelCapacity.ReachOn ๐ P u v โ HiddenChannelCapacity.ReachOn ๐ P v w โ HiddenChannelCapacity.ReachOn ๐ P u wโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v : ฮน},
HiddenChannelCapacity.ReachOn ๐ P u v โ HiddenChannelCapacity.ReachOn ๐ P v uโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u : ฮน}, P u โ HiddenChannelCapacity.ReachOn ๐ P u uEvery vertex on a walk is reachable from its head.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v : ฮน} {p : List ฮน},
HiddenChannelCapacity.IsWalkOn ๐ P p โ p.head? = some u โ v โ p โ HiddenChannelCapacity.ReachOn ๐ P u vโ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} {u v w : ฮน},
HiddenChannelCapacity.ReachOn ๐ P u v โ HiddenChannelCapacity.Adj ๐ v w โ P w โ HiddenChannelCapacity.ReachOn ๐ P u wโ {ฮน : Type u_1} {Q : List ฮน โ Prop}, (โ p, Q p) โ โ p, Q p โง โ (q : List ฮน), Q q โ p.length โค q.lengthA connector: an x-y path with interior inside Q, from adjacent entry points and interior reachability.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {Q : ฮน โ Prop} {x y uโ uโ : ฮน},
x โ y โ
HiddenChannelCapacity.Adj ๐ x uโ โ
HiddenChannelCapacity.Adj ๐ y uโ โ
HiddenChannelCapacity.ReachOn ๐ Q uโ uโ โ
ยฌQ x โ
ยฌQ y โ
โ p,
List.IsChain (HiddenChannelCapacity.Adj ๐) p โง
p.Nodup โง p.head? = some x โง p.getLast? = some y โง โ v โ p, v = x โจ v = y โจ Q vIf x has no neighbor reachable from a off the separator, then erasing x still separates: the first x-occurrence on any violating path has its predecessor reachable from a and adjacent to x.
โ {ฮน : Type u_1} [inst : DecidableEq ฮน] {๐ : Finset (Finset ฮน)} {S Sep : Finset ฮน} {a b x : ฮน},
x โ a โ
(โ (u : ฮน),
u โ S \ Sep โง HiddenChannelCapacity.ReachOn ๐ (fun v => v โ S \ Sep) a u โ ยฌHiddenChannelCapacity.Adj ๐ x u) โ
ยฌHiddenChannelCapacity.ReachOn ๐ (fun v => v โ S \ Sep) a b โ
ยฌHiddenChannelCapacity.ReachOn ๐ (fun v => v โ S \ Sep.erase x) a bEvery chained walk contains a duplicate-free path with the same endpoints and vertices.
โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {P : ฮน โ Prop} (p : List ฮน),
HiddenChannelCapacity.IsWalkOn ๐ P p โ
โ {u v : ฮน}, p.head? = some u โ p.getLast? = some v โ HiddenChannelCapacity.ReachOn ๐ P u vโ {ฮน : Type u_1} {R : ฮน โ ฮน โ Prop} {p : List ฮน}, List.IsChain R p โ โ (n : โ), List.IsChain R (List.drop n p)โ {ฮน : Type u_1} {R : ฮน โ ฮน โ Prop} {p : List ฮน}, List.IsChain R p โ โ (n : โ), List.IsChain R (List.take n p)Positional form of the head.
โ {ฮน : Type u_1} {p : List ฮน} {x : ฮน}, p.head? = some x โ โ (h0 : 0 < p.length), p[0] = xโ {ฮน : Type u_1} (l : List ฮน) {s t : โ} (hs : s < l.length) (ht : t < l.length), s = t โ l[s] = l[t]โ {ฮน : Type u_1} {๐ : Finset (Finset ฮน)} {u v : ฮน}, HiddenChannelCapacity.Adj ๐ u v โ HiddenChannelCapacity.Adj ๐ v uยฌโ l, HiddenChannelCapacity.IsChordlessCycle ๐ l
Primal adjacency: distinct co-hosted vertices.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ ฮน โ ฮน โ PropA chordless cycle: duplicate-free, length โฅ 4, consecutive pairs hosted, and the only adjacencies among its vertices are the cyclically consecutive ones.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ Finset (Finset ฮน) โ List ฮน โ PropA walk constrained to P: chained adjacencies, every vertex satisfies P.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ (ฮน โ Prop) โ List ฮน โ PropReachability within P: a duplicate-free walk with the given endpoints.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ (ฮน โ Prop) โ ฮน โ ฮน โ Propv is simplicial within the vertex set S: it lies in S and its ๐-neighbors inside S are pairwise adjacent.
{ฮน : Type u_1} โ Finset (Finset ฮน) โ Finset ฮน โ ฮน โ PropThe consecutive-pair supports of a cyclic vertex list.
{ฮน : Type u_1} โ [DecidableEq ฮน] โ List ฮน โ List (Finset ฮน){ฮน : Type u_1} โ [DecidableEq ฮน] โ (๐ : Finset (Finset ฮน)) โ (u v : ฮน) โ Decidable (HiddenChannelCapacity.Adj ๐ u v)The two-clique quantum junction tree theorem (LauritzenโZwiernik, arXiv:2605.19453, Theorem 3.1): over the acyclic two-clique base, the base's compatibility datum decides base-side reconstruction. For strictly positive consistent marginals, a quantum Markov completion exists iff Tr T(R) = 1; the completion is then unique and equals the normalized logarithmic candidate ฯ(R) = T(R)/Tr T(R). Together with trC_le_one (the trace bound) this is the full theorem.
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b]
[inst_8 : Nonempty b] {ฯAC : MState (a ร c)} {ฯCB : MState (c ร b)},
ฯAC.m.PosDef โ
ฯCB.m.PosDef โ
ฯAC.traceLeft = ฯCB.traceRight โ
(HiddenChannelCapacity.trC ฯAC ฯCB = 1 โ โ ฯ, HiddenChannelCapacity.IsMarkovCompletion ฯAC ฯCB ฯ) โง
โ (ฯ : MState (a ร c ร b)),
HiddenChannelCapacity.IsMarkovCompletion ฯAC ฯCB ฯ โ ฯ = HiddenChannelCapacity.sigT ฯAC ฯCBโ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b]
[inst_8 : Nonempty b] {ฯAC : MState (a ร c)} {ฯCB : MState (c ร b)},
ฯAC.m.PosDef โ
ฯCB.m.PosDef โ
ฯAC.traceLeft = ฯCB.traceRight โ
(HiddenChannelCapacity.trC ฯAC ฯCB = 1 โ โ ฯ, HiddenChannelCapacity.IsMarkovCompletion ฯAC ฯCB ฯ) โง
โ (ฯ : MState (a ร c ร b)),
HiddenChannelCapacity.IsMarkovCompletion ฯAC ฯCB ฯ โ ฯ = HiddenChannelCapacity.sigT ฯAC ฯCBLZ Theorem 3.1, existence direction: if Tr T(R) = 1 then the normalized candidate IS a quantum Markov completion (extraction of ฮ_R = 0 via DPI and the faithfulness of relative entropy, in the order recorded in the QJT_URS front-load).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b]
[inst_8 : Nonempty b] {ฯAC : MState (a ร c)} {ฯCB : MState (c ร b)},
ฯAC.m.PosDef โ
ฯCB.m.PosDef โ
ฯAC.traceLeft = ฯCB.traceRight โ
HiddenChannelCapacity.trC ฯAC ฯCB = 1 โ
HiddenChannelCapacity.IsMarkovCompletion ฯAC ฯCB (HiddenChannelCapacity.sigT ฯAC ฯCB)LZ Theorem 3.1, uniqueness direction: any quantum Markov completion forces Tr T(R) = 1 and equals the normalized candidate ฯ(R).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b]
[inst_8 : Nonempty b] {ฯAC : MState (a ร c)} {ฯCB : MState (c ร b)},
ฯAC.m.PosDef โ
ฯCB.m.PosDef โ
ฯAC.traceLeft = ฯCB.traceRight โ
โ (ฯ : MState (a ร c ร b)),
HiddenChannelCapacity.IsMarkovCompletion ฯAC ฯCB ฯ โ
HiddenChannelCapacity.trC ฯAC ฯCB = 1 โง ฯ = HiddenChannelCapacity.sigT ฯAC ฯCBLZ Theorem 3.1, the trace bound: Tr T(R) โค 1. Proved WITHOUT Lieb's three-matrix inequality, by the eq. 13 route specialized to two cliques (Lemma 3.2 at ฯ = ฯ(R) plus SSA and DPI); see the QJT_URS L2a front-load.
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b] [Nonempty b]
{ฯAC : MState (a ร c)} {ฯCB : MState (c ร b)}, ฯAC.m.PosDef โ ฯCB.m.PosDef โ HiddenChannelCapacity.trC ฯAC ฯCB โค 1ฮ_R(ฯ) โฅ 0: one DPI application plus Klein nonnegativity (LZ Lem 3.2's nonnegativity clause, via Prop A.7 imported through the vendored channel DPI).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Fintype b] [inst_6 : DecidableEq b] {ฯAC : MState (a ร c)}
{ฯCB : MState (c ร b)},
ฯAC.m.PosDef โ
ฯCB.m.PosDef โ
โ (ฯ : MState (a ร c ร b)),
0 โค
HiddenChannelCapacity.qDivR (HiddenChannelCapacity.margAC ฯ) ฯAC +
HiddenChannelCapacity.qDivR (HiddenChannelCapacity.margCB ฯ) ฯCB -
HiddenChannelCapacity.qDivR (HiddenChannelCapacity.margC ฯ) ฯAC.traceLeftDPI for the real divergence under the left partial trace (monotonicity of relative entropy, LZ Prop A.7; imported from the vendored channel DPI sandwichedRenyiEntropy_DPI_eq_one at the partial-trace channel).
โ {dโ : Type u_5} {dโ : Type u_6} [inst : Fintype dโ] [inst_1 : DecidableEq dโ] [inst_2 : Fintype dโ]
[inst_3 : DecidableEq dโ] (ฯ ฯ : MState (dโ ร dโ)),
ฯ.m.PosDef โ
ฯ.traceLeft.m.PosDef โ HiddenChannelCapacity.qDivR ฯ.traceLeft ฯ.traceLeft โค HiddenChannelCapacity.qDivR ฯ ฯThe single ENNReal-to-โ seam: qDivR equals the vendored relative entropy. (LZ p. 5; vendored qRelativeEnt_rank.)
โ {d : Type u_4} [inst : Fintype d] [inst_1 : DecidableEq d] {ฯ ฯ : MState d},
ฯ.m.PosDef โ HiddenChannelCapacity.qDivR ฯ ฯ = ๐(ฯโฯ).toRealโ {x : ENNReal} {r : โ}, x โ โค โ โx = โr โ x.toReal = rThe partial trace of a strictly positive state is strictly positive (marginals of elements of Sโโบ stay in Sโโบ, LZ p. 4).
โ {dโ : Type u_1} {dโ : Type u_2} [inst : Fintype dโ] [inst_1 : DecidableEq dโ] [Nonempty dโ] [inst_3 : Fintype dโ]
[inst_4 : DecidableEq dโ] {ฯ : MState (dโ ร dโ)}, ฯ.m.PosDef โ ฯ.traceLeft.m.PosDefโ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b]
[inst_8 : Nonempty b] (ฯAC : MState (a ร c)) (ฯCB : MState (c ร b)), (HiddenChannelCapacity.sigT ฯAC ฯCB).m.PosDefโ {d : Type u_4} [inst : Fintype d] [inst_1 : DecidableEq d] (ฯ : MState d), HiddenChannelCapacity.qDivR ฯ ฯ = 0Klein nonnegativity in real form (vendored core: inner_log_sub_log_nonneg).
โ {d : Type u_4} [inst : Fintype d] [inst_1 : DecidableEq d] {ฯ ฯ : MState d},
ฯ.m.PosDef โ 0 โค HiddenChannelCapacity.qDivR ฯ ฯโ {d : Type u_4} [inst : Fintype d] [inst_1 : DecidableEq d] {ฯ ฯ : MState d}, ฯ.m.PosDef โ (โฯ).ker โค (โฯ).kerLZ Lemma 3.2, normalized form: for EVERY state ฯ, D(ฯโฯ(R)) โ log Tr T(R) = I(A:B|C)_ฯ + ฮ_R(ฯ) with ฮ_R(ฯ) := D(ฯ_ACโฯ_AC) + D(ฯ_CBโฯ_CB) โ D(ฯ_Cโฯ_C). This is the paper's divergence identity (p. 8) restated against the NORMALIZED candidate via log T(R) = log Tr T(R) + log ฯ(R) (the presentation variant recorded in QJT_URS; pure real algebra over the duality, no positivity hypotheses).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b]
[inst_8 : Nonempty b] (ฯAC : MState (a ร c)) (ฯCB : MState (c ร b)) (ฯ : MState (a ร c ร b)),
HiddenChannelCapacity.qDivR ฯ (HiddenChannelCapacity.sigT ฯAC ฯCB) - Real.log (HiddenChannelCapacity.trC ฯAC ฯCB) =
qcmi ฯ +
(HiddenChannelCapacity.qDivR (HiddenChannelCapacity.margAC ฯ) ฯAC +
HiddenChannelCapacity.qDivR (HiddenChannelCapacity.margCB ฯ) ฯCB -
HiddenChannelCapacity.qDivR (HiddenChannelCapacity.margC ฯ) ฯAC.traceLeft)The log of the normalized candidate splits into the trace constant and the embedded-log sum (log T(R) = log Tr T(R) + log ฯ(R), the normalization step of the LZ eq. 13 route).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [inst_5 : Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b]
[inst_8 : Nonempty b] (ฯAC : MState (a ร c)) (ฯCB : MState (c ร b)),
(โ(HiddenChannelCapacity.sigT ฯAC ฯCB)).log =
Real.log (HiddenChannelCapacity.trC ฯAC ฯCB)โปยน โข 1 + HiddenChannelCapacity.Tmat ฯAC ฯCBโ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [Nonempty a]
[inst_3 : Fintype c] [inst_4 : DecidableEq c] [Nonempty c] [inst_6 : Fintype b] [inst_7 : DecidableEq b] [Nonempty b]
(ฯAC : MState (a ร c)) (ฯCB : MState (c ร b)), 0 < HiddenChannelCapacity.trC ฯAC ฯCBexp then log is the identity on Hermitian matrices (unconditionally: both are cfc transports and Real.log โ Real.exp = id).
โ {d : Type u_4} [inst : Fintype d] [inst_1 : DecidableEq d] (A : HermitianMat d โ), A.exp.log = ADuality for the C-embedding (two-step pull-out; uses the marginal coherence margC_eq_CBright, LZ Lem 2.2).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Fintype c]
[inst_3 : DecidableEq c] [inst_4 : Fintype b] [inst_5 : DecidableEq b] (ฯ : MState (a ร c ร b))
(L : HermitianMat c โ), inner โ (โฯ) (HiddenChannelCapacity.embedC L) = inner โ (โ(HiddenChannelCapacity.margC ฯ)) LCoherence of the two separator-marginal routes (iterated marginalization, LZ Lem 2.2): tracing a then b is tracing b then a.
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Fintype c]
[inst_3 : DecidableEq c] [inst_4 : Fintype b] [inst_5 : DecidableEq b] (ฯ : MState (a ร c ร b)),
HiddenChannelCapacity.margC ฯ = (HiddenChannelCapacity.margCB ฯ).traceRightDuality for a right-identity Kronecker factor on any bipartite carrier.
โ {dโ : Type u_4} {dโ : Type u_5} [inst : Fintype dโ] [inst_1 : DecidableEq dโ] [inst_2 : Fintype dโ]
[inst_3 : DecidableEq dโ] (X : MState (dโ ร dโ)) (L : HermitianMat dโ โ),
inner โ (โX) (L.kronecker 1) = inner โ (โX.traceRight) LPull-out for the right partial trace, right-multiplication form: Tr_B(ฯ (M โ 1_B)) = Tr_B(ฯ) ยท M.
โ {R : Type u_1} [inst : NonAssocSemiring R] {d : Type u_2} {dโ : Type u_3} {dโ : Type u_4} {dโ : Type u_5}
[inst_1 : Fintype d] [inst_2 : Fintype dโ] [inst_3 : DecidableEq d] (ฯ : Matrix (dโ ร d) (dโ ร d) R)
(M : Matrix dโ dโ R), Matrix.traceRight (ฯ * Matrix.kroneckerMap (fun x1 x2 => x1 * x2) M 1) = ฯ.traceRight * MDuality for the CโชB-embedding: testing the state against an embedded observable is testing the marginal (LZ eq. 2).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Fintype c]
[inst_3 : DecidableEq c] [inst_4 : Fintype b] [inst_5 : DecidableEq b] (ฯ : MState (a ร c ร b))
(L : HermitianMat (c ร b) โ),
inner โ (โฯ) (HiddenChannelCapacity.embedCB L) = inner โ (โ(HiddenChannelCapacity.margCB ฯ)) LPull-out, right-multiplication form: Tr_A(ฯ (1_A โ M)) = Tr_A(ฯ) ยท M.
โ {R : Type u_1} [inst : NonAssocSemiring R] {d : Type u_2} {dโ : Type u_3} {dโ : Type u_4} {dโ : Type u_5}
[inst_1 : Fintype d] [inst_2 : Fintype dโ] [inst_3 : DecidableEq d] (ฯ : Matrix (d ร dโ) (d ร dโ) R)
(M : Matrix dโ dโ R), Matrix.traceLeft (ฯ * Matrix.kroneckerMap (fun x1 x2 => x1 * x2) 1 M) = ฯ.traceLeft * MDuality for the AโชC-embedding (via the L0 primitive traceAlong_duality).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Fintype c]
[inst_3 : DecidableEq c] [inst_4 : Fintype b] [inst_5 : DecidableEq b] (ฯ : MState (a ร c ร b))
(L : HermitianMat (a ร c) โ),
inner โ (โฯ) (HiddenChannelCapacity.embedAC L) = inner โ (โ(HiddenChannelCapacity.margAC ฯ)) LThe partial-trace duality (LZ eq (2), the defining property of the partial trace, split-indexed): testing the marginal against an observable is testing the state against the embedded observable, Tr(embedAlong e M ยท ฯ) = Tr(M ยท traceAlong e ฯ).
โ {d : Type u_1} {a : Type u_2} {b : Type u_3} [inst : Fintype d] [inst_1 : DecidableEq d] [inst_2 : Fintype a]
[inst_3 : DecidableEq a] [inst_4 : Fintype b] [inst_5 : DecidableEq b] (ฯ : MState d) (e : d โ a ร b)
(M : Matrix a a โ), (Matrix.embedAlong e M * ฯ.m).trace = (M * (ฯ.traceAlong e).m).traceThe trace is invariant under simultaneous reindexing by an equivalence.
โ {R : Type u_1} [inst : NonAssocSemiring R] {n : Type u_6} {m : Type u_7} [inst_1 : Fintype n] [inst_2 : Fintype m]
(A : Matrix m m R) (e : n โ m), (A.submatrix โe โe).trace = A.tracePull-out for the right partial trace, left-multiplication form: Tr_B((M โ 1_B) ฯ) = M ยท Tr_B(ฯ).
โ {R : Type u_1} [inst : NonAssocSemiring R] {d : Type u_2} {dโ : Type u_3} {dโ : Type u_4} {dโ : Type u_5}
[inst_1 : Fintype d] [inst_2 : Fintype dโ] [inst_3 : DecidableEq d] (M : Matrix dโ dโ R)
(ฯ : Matrix (dโ ร d) (dโ ร d) R),
Matrix.traceRight (Matrix.kroneckerMap (fun x1 x2 => x1 * x2) M 1 * ฯ) = M * ฯ.traceRightThe AC-marginal in traceAlong form (the L0 primitive).
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Fintype c]
[inst_3 : DecidableEq c] [inst_4 : Fintype b] [inst_5 : DecidableEq b] (ฯ : MState (a ร c ร b)),
HiddenChannelCapacity.margAC ฯ = ฯ.traceAlong (Equiv.prodAssoc a c b).symmThe vendored associator is the plain prodAssoc relabel.
โ {a : Type u_1} {c : Type u_2} {b : Type u_3} [inst : Fintype a] [inst_1 : DecidableEq a] [inst_2 : Fintype c]
[inst_3 : DecidableEq c] [inst_4 : Fintype b] [inst_5 : DecidableEq b] (ฯ : MState (a ร c ร b)),
ฯ.assoc' = ฯ.relabel (Equiv.prodAssoc a c b)Faithfulness of quantum relative entropy (the equality case of Klein's inequality, LZ p. 5, Ruskai 2002 Thm 3; not present in the vendored corpus): for a state ฯ and a STRICTLY POSITIVE state ฯ, Tr[ฯ(log ฯ โ log ฯ)] = 0 forces ฯ = ฯ.
โ {d : Type u_1} [inst : Fintype d] [inst_1 : DecidableEq d] (ฯ ฯ : MState d),
ฯ.m.PosDef โ inner โ (โฯ) ((โฯ).log - (โฯ).log) = 0 โ ฯ = ฯโ {x y : โ}, 0 โค x โ 0 < y โ 0 โค HiddenChannelCapacity.kleinTerm x yโ {x y : โ}, 0 โค x โ 0 < y โ (HiddenChannelCapacity.kleinTerm x y = 0 โ x = y)Rows of the shared-basis kernel sum to one (C is unitary).
โ {d : Type u_1} [inst : Fintype d] [inst_1 : DecidableEq d] (A B : HermitianMat d โ) (i : d),
โ j, โ((โโฏ.eigenvectorUnitary).conjTranspose * โโฏ.eigenvectorUnitary) i jโ ^ 2 = 1Columns of the shared-basis kernel sum to one.
โ {d : Type u_1} [inst : Fintype d] [inst_1 : DecidableEq d] (A B : HermitianMat d โ) (j : d),
โ i, โ((โโฏ.eigenvectorUnitary).conjTranspose * โโฏ.eigenvectorUnitary) i jโ ^ 2 = 1A.log is cfc at Real.log, definitionally (QuantumInfo LogExp.lean).
โ {d : Type u_1} [inst : Fintype d] [inst_1 : DecidableEq d] (A : HermitianMat d โ), A.log = A.cfc Real.logThe shared-basis double sum (generalizing the vendored inner_eq_doubly_stochastic_sum from (A, B) to (A.cfc f, B.cfc g) with the SAME overlap kernel): the inner product of two matrix functions expands over the pair of eigenbases with weights โC i jโยฒ, C = U_Aโ U_B. All four divergence pairings of the faithfulness argument are instances of this one lemma at a single shared kernel.
โ {d : Type u_1} [inst : Fintype d] [inst_1 : DecidableEq d] (A B : HermitianMat d โ) (f g : โ โ โ),
inner โ (A.cfc f) (B.cfc g) =
โ i,
โ j,
f (โฏ.eigenvalues i) * g (โฏ.eigenvalues j) *
โ((โโฏ.eigenvectorUnitary).conjTranspose * โโฏ.eigenvectorUnitary) i jโ ^ 2โ {d : Type u_1} [inst : Fintype d] [inst_1 : DecidableEq d] (C : Matrix d d โ) (i j : d),
(Matrix.single i i 1 * C * Matrix.single j j 1 * C.conjTranspose).trace = C i j * star (C i j)โ {d : Type u_1} [inst : Fintype d] [inst_1 : DecidableEq d] (X : Matrix d d โ) (i : d),
(Matrix.single i i 1 * X).trace = X i iฯAC.m.PosDef
ฯCB.m.PosDef
ฯAC.traceLeft = ฯCB.traceRight
A completion: a global state with the prescribed marginals (LZ M(R), p. 7).
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ
[inst_4 : Fintype b] โ
[inst_5 : DecidableEq b] โ MState (a ร c) โ MState (c ร b) โ MState (a ร c ร b) โ PropA quantum Markov completion: a completion with I(A:B|C) = 0 (LZ Def 2.4 quantum conditional independence + the Markov completion of p. 7).
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ
[inst_4 : Fintype b] โ
[inst_5 : DecidableEq b] โ MState (a ร c) โ MState (c ร b) โ MState (a ร c ร b) โ PropThe logarithmic candidate T(R) = exp(log ฯ_{AC} + log ฯ_{CB} โ log ฯ_C) (LZ eq. 5). A Hermitian matrix, NOT a priori a state: its trace deficit is the subject of the theorem.
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ
[inst_4 : Fintype b] โ
[inst_5 : DecidableEq b] โ MState (a ร c) โ MState (c ร b) โ HermitianMat (a ร c ร b) โThe Hermitian exponent of the logarithmic candidate: the sum of the embedded logarithms of the marginals (log ฯ_{AC} + log ฯ_{CB} โ log ฯ_C, every operator embedded BEFORE exp; LZ eq. 5 with the p. 23 embedding convention; the separator marginal is taken from the AC side).
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ
[inst_4 : Fintype b] โ
[inst_5 : DecidableEq b] โ MState (a ร c) โ MState (c ร b) โ HermitianMat (a ร c ร b) โEmbed an AโชC-local operator into the global system.
{a : Type u_1} โ {c : Type u_2} โ {b : Type u_3} โ [DecidableEq b] โ HermitianMat (a ร c) โ โ HermitianMat (a ร c ร b) โEmbed a C-local operator into the global system.
{a : Type u_1} โ
{c : Type u_2} โ {b : Type u_3} โ [DecidableEq a] โ [DecidableEq b] โ HermitianMat c โ โ HermitianMat (a ร c ร b) โEmbed a CโชB-local operator into the global system (tensoring with the identity, LZ p. 23).
{a : Type u_1} โ {c : Type u_2} โ {b : Type u_3} โ [DecidableEq a] โ HermitianMat (c ร b) โ โ HermitianMat (a ร c ร b) โThe scalar Klein term x log x โ x log y โ x + y (the integrand of Klein's inequality, LZ p. 5 / Ruskai 2002 Thm 3).
โ โ โ โ โ
The AโชC-marginal of a global state (LZ p. 4: reduced density operator).
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ
[inst_4 : Fintype b] โ [inst_5 : DecidableEq b] โ MState (a ร c ร b) โ MState (a ร c)The C-marginal (the separator marginal), as the vendor composite the qcmi definition consumes.
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ [inst_4 : Fintype b] โ [inst_5 : DecidableEq b] โ MState (a ร c ร b) โ MState cThe CโชB-marginal of a global state.
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ
[inst_4 : Fintype b] โ [inst_5 : DecidableEq b] โ MState (a ร c ร b) โ MState (c ร b)The quantum relative entropy in real form, Tr[ฯ (log ฯ โ log ฯ)] (LZ eq. 4 on density operators; equals the vendored qRelativeEnt by qRelativeEnt_rank when ฯ is nonsingular).
{d : Type u_4} โ [inst : Fintype d] โ [inst_1 : DecidableEq d] โ MState d โ MState d โ โThe normalized candidate ฯ(R) = T(R)/Tr T(R): the state the equivalences of LZ Thm 3.1 are about.
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[Nonempty a] โ
[inst_3 : Fintype c] โ
[inst_4 : DecidableEq c] โ
[Nonempty c] โ
[inst_6 : Fintype b] โ
[inst_7 : DecidableEq b] โ [Nonempty b] โ MState (a ร c) โ MState (c ร b) โ MState (a ร c ร b)The trace of the logarithmic candidate.
{a : Type u_1} โ
{c : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype a] โ
[inst_1 : DecidableEq a] โ
[inst_2 : Fintype c] โ
[inst_3 : DecidableEq c] โ
[inst_4 : Fintype b] โ [inst_5 : DecidableEq b] โ MState (a ร c) โ MState (c ร b) โ โTrace out the second block of a coordinate split: the abstract partial-trace primitive of the L0 layer.
{d : Type u_1} โ
{a : Type u_2} โ
{b : Type u_3} โ
[inst : Fintype d] โ
[inst_1 : DecidableEq d] โ
[inst_2 : Fintype a] โ
[inst_3 : DecidableEq a] โ [Fintype b] โ [DecidableEq b] โ MState d โ d โ a ร b โ MState aEmbed an operator on one block of a coordinate split into the whole space (tensor with the identity, transported along the split): the dual of MState.traceAlong, and the paper's embedding convention ("tensoring with identities BEFORE logs") as a named primitive.
{R : Type u_1} โ
[NonAssocSemiring R] โ
{d : Type u_2} โ {a : Type u_6} โ {b : Type u_7} โ [DecidableEq b] โ d โ a ร b โ Matrix a a R โ Matrix d d RClosed form of the multivariate Gaussian KL divergence (M3b). For positive-definite covariances Sโ, Sโ, KL(N(mโ,Sโ) โ N(mโ,Sโ)) = ยฝ ( log(det Sโ / det Sโ) + tr(Sโโปยน Sโ) + โชmโ-mโ, Sโโปยน(mโ-mโ)โซ - d ).
The whitening reduction (klDivReal_multivariateGaussian_whiten) sends the pair to a KL against the standard Gaussian (klDivReal_multivariateGaussian_stdGaussian); the three scalar invariants of the whitened covariance (det_whitened, trace_whitened, normSq_cfcSqrt_inv_apply) then identify the arguments.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (mโ mโ : EuclideanSpace โ ฮน) {Sโ Sโ : Matrix ฮน ฮน โ},
Sโ.PosDef โ
Sโ.PosDef โ
InformationTheory.klDivReal (ProbabilityTheory.multivariateGaussian mโ Sโ)
(ProbabilityTheory.multivariateGaussian mโ Sโ) =
1 / 2 *
(Real.log (Sโ.det / Sโ.det) + (Sโโปยน * Sโ).trace + (mโ - mโ).ofLp โฌแตฅ Sโโปยน.mulVec (mโ.ofLp - mโ.ofLp) -
โ(Fintype.card ฮน))โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (mโ mโ : EuclideanSpace โ ฮน) {Sโ Sโ : Matrix ฮน ฮน โ},
Sโ.PosDef โ
Sโ.PosDef โ
InformationTheory.klDivReal (ProbabilityTheory.multivariateGaussian mโ Sโ)
(ProbabilityTheory.multivariateGaussian mโ Sโ) =
1 / 2 *
(Real.log (Sโ.det / Sโ.det) + (Sโโปยน * Sโ).trace + (mโ - mโ).ofLp โฌแตฅ Sโโปยน.mulVec (mโ.ofLp - mโ.ofLp) -
โ(Fintype.card ฮน))tr C(Sโ,Sโ) = tr(Sโโปยน Sโ).
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] {Sโ Sโ : Matrix ฮน ฮน โ},
Sโ.PosDef โ
Sโ.PosDef โ
((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ * ((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ).conjTranspose).trace = (Sโโปยน * Sโ).traceThe whitened covariance (โSโโปยน โSโ)(โSโโปยน โSโ)แดด is positive-definite (its factor is invertible).
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] {Sโ Sโ : Matrix ฮน ฮน โ},
Sโ.PosDef โ Sโ.PosDef โ ((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ * ((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ).conjTranspose).PosDefThe whitening map preserves the quadratic form: โโSโโปยน vโยฒ = โชv, Sโโปยน vโซ.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] {Sโ : Matrix ฮน ฮน โ},
Sโ.PosDef โ โ (v : EuclideanSpace โ ฮน), โ(Matrix.toEuclideanCLM (CFC.sqrt Sโ)โปยน) vโ ^ 2 = v.ofLp โฌแตฅ Sโโปยน.mulVec v.ofLpThe inverse of a functional-calculus square root is self-adjoint.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (S : Matrix ฮน ฮน โ),
(CFC.sqrt S)โปยน.conjTranspose = (CFC.sqrt S)โปยนThe functional-calculus square root of a matrix is self-adjoint.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (S : Matrix ฮน ฮน โ), (CFC.sqrt S).conjTranspose = CFC.sqrt S(โS)โปยน (โS)โปยน = Sโปยน for positive-definite S.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] {S : Matrix ฮน ฮน โ},
S.PosDef โ (CFC.sqrt S)โปยน * (CFC.sqrt S)โปยน = SโปยนWhitening reduction of the multivariate Gaussian KL. For positive-definite Sโ, pushing both measures through the whitening equivalence (gaussianAffineEquiv mโ hSโ).symm reduces the KL divergence of two multivariate Gaussians to a KL against the standard Gaussian, with whitened mean โSโโปยน (mโ โ mโ) and covariance (โSโโปยน โSโ)(โSโโปยน โSโ)แดด.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (mโ mโ : EuclideanSpace โ ฮน) (Sโ Sโ : Matrix ฮน ฮน โ),
Sโ.PosDef โ
InformationTheory.klDivReal (ProbabilityTheory.multivariateGaussian mโ Sโ)
(ProbabilityTheory.multivariateGaussian mโ Sโ) =
InformationTheory.klDivReal
(ProbabilityTheory.multivariateGaussian ((Matrix.toEuclideanCLM (CFC.sqrt Sโ)โปยน) (mโ - mโ))
((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ * ((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ).conjTranspose))
(ProbabilityTheory.stdGaussian (EuclideanSpace โ ฮน))multivariateGaussian m S is the pushforward of the standard Gaussian by the affine equivalence gaussianAffineEquiv m hS.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (m : EuclideanSpace โ ฮน) {S : Matrix ฮน ฮน โ}
(hS : S.PosDef),
ProbabilityTheory.multivariateGaussian m S =
MeasureTheory.Measure.map (โ(ProbabilityTheory.gaussianAffineEquiv m hS))
(ProbabilityTheory.stdGaussian (EuclideanSpace โ ฮน))gaussianAffineEquiv m hS is the affine map z โฆ m + โS z.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (m : EuclideanSpace โ ฮน) {S : Matrix ฮน ฮน โ} (hS : S.PosDef)
(z : EuclideanSpace โ ฮน), (ProbabilityTheory.gaussianAffineEquiv m hS) z = m + (Matrix.toEuclideanCLM (CFC.sqrt S)) zThe whitening map: the inverse of gaussianAffineEquiv m hS is z โฆ โSโปยน (z - m).
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (m : EuclideanSpace โ ฮน) {S : Matrix ฮน ฮน โ} (hS : S.PosDef)
(z : EuclideanSpace โ ฮน),
(ProbabilityTheory.gaussianAffineEquiv m hS).symm z = (Matrix.toEuclideanCLM (CFC.sqrt S)โปยน) (z - m)Closed form of the multivariate Gaussian KL against the standard Gaussian. For a positive-definite covariance C, klDivReal (N(w, C)) (stdGaussian) = ยฝ ( -log det C + tr C + โwโยฒ - d ).
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (w : EuclideanSpace โ ฮน) {C : Matrix ฮน ฮน โ},
C.PosDef โ
InformationTheory.klDivReal (ProbabilityTheory.multivariateGaussian w C)
(ProbabilityTheory.stdGaussian (EuclideanSpace โ ฮน)) =
1 / 2 * (-Real.log C.det + C.trace + โwโ ^ 2 - โ(Fintype.card ฮน))An orthogonal change of variables preserves the Euclidean norm: if Mแดด M = 1 then โtoEuclideanCLM M wโยฒ = โwโยฒ.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] {M : Matrix ฮน ฮน โ},
M.conjTranspose * M = 1 โ โ (w : EuclideanSpace โ ฮน), โ(Matrix.toEuclideanCLM M) wโ ^ 2 = โwโ ^ 2General affine image of a multivariate Gaussian. For positive-semidefinite covariance S, the affine pushforward x โฆ a + B x of multivariateGaussian ฮผ S is the multivariate Gaussian with transported mean a + B ฮผ and congruent covariance B * S * Bแดด. This is the covariance- transport engine of the diagonalization step: an orthogonal B = Uแต sends the whitened covariance to its eigenvalue diagonal Uแต S U. Proved by composing the two affine maps and reducing to the standard-Gaussian image map_stdGaussian_affine, using โS ยท โSแดด = โS ยท โS = S.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (a ฮผ : EuclideanSpace โ ฮน) (B S : Matrix ฮน ฮน โ),
S.PosSemidef โ
MeasureTheory.Measure.map (fun x => a + (Matrix.toEuclideanCLM B) x) (ProbabilityTheory.multivariateGaussian ฮผ S) =
ProbabilityTheory.multivariateGaussian (a + (Matrix.toEuclideanCLM B) ฮผ) (B * S * B.conjTranspose)KL of a diagonal multivariate Gaussian against the standard Gaussian, tensorized. For strictly positive diagonal covariance d > 0, klDivReal (mvG ฮผ (diagonal d)) stdGaussian = โแตข ยฝ(-log dแตข + dแตข + ฮผแตขยฒ - 1), the coordinatewise sum of the scalar Gaussian-Gaussian KL closed forms. This is the diagonalized-and-tensorized leg of the multivariate Gaussian KL closed form.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (ฮผ : EuclideanSpace โ ฮน) (d : ฮน โ โ),
(โ (i : ฮน), 0 < d i) โ
InformationTheory.klDivReal (ProbabilityTheory.multivariateGaussian ฮผ (Matrix.diagonal d))
(ProbabilityTheory.stdGaussian (EuclideanSpace โ ฮน)) =
โ i, 1 / 2 * (-Real.log (d i) + d i + ฮผ.ofLp i ^ 2 - 1)A diagonal multivariate Gaussian is a pushed-forward product of scalar Gaussians. For d โฅ 0, multivariateGaussian ฮผ (diagonal d) is the image under toLp 2 of the product measure โแตข gaussianReal ฮผแตข dแตข.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (ฮผ : EuclideanSpace โ ฮน) (d : ฮน โ โ),
(โ (i : ฮน), 0 โค d i) โ
ProbabilityTheory.multivariateGaussian ฮผ (Matrix.diagonal d) =
MeasureTheory.Measure.map (WithLp.toLp 2)
(MeasureTheory.Measure.pi fun i => ProbabilityTheory.gaussianReal (ฮผ.ofLp i) (d i).toNNReal)The affine image z โฆ a + B z of the standard Gaussian on EuclideanSpace โ ฮน is the multivariate Gaussian with mean a and covariance B * Bแดด.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (a : EuclideanSpace โ ฮน) (B : Matrix ฮน ฮน โ),
MeasureTheory.Measure.map (fun z => a + (Matrix.toEuclideanCLM B) z)
(ProbabilityTheory.stdGaussian (EuclideanSpace โ ฮน)) =
ProbabilityTheory.multivariateGaussian a (B * B.conjTranspose)The real inner product of the two adjoint images (toEuclideanCLM B)แดด x and (toEuclideanCLM B)แดด y equals the quadratic form x โฌแตฅ (B * Bแดด) *แตฅ y.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (B : Matrix ฮน ฮน โ) (x y : EuclideanSpace โ ฮน),
inner โ ((ContinuousLinearMap.adjoint (Matrix.toEuclideanCLM B)) x)
((ContinuousLinearMap.adjoint (Matrix.toEuclideanCLM B)) y) =
x.ofLp โฌแตฅ (B * B.conjTranspose).mulVec y.ofLpThe adjoint of toEuclideanCLM B is toEuclideanCLM Bแดด: toEuclideanCLM is a star algebra equivalence, so it carries the matrix conjugate transpose to the operator adjoint.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (B : Matrix ฮน ฮน โ),
ContinuousLinearMap.adjoint (Matrix.toEuclideanCLM B) = Matrix.toEuclideanCLM B.conjTransposeFiniteness of the one-dimensional Gaussian-Gaussian KL divergence. For variance parameters vโ โ 0, vโ โ 0, the โโฅ0โ-valued klDiv (N(mโ, vโ)) (N(mโ, vโ)) is finite. This is the integrability side-condition that klDivReal_pi requires per coordinate: the log-likelihood ratio is P-a.e. an explicit quadratic, whose square-integrability against P = N(mโ, vโ) is the finiteness of the second moment.
โ {mโ mโ : โ} {vโ vโ : NNReal},
vโ โ 0 โ
vโ โ 0 โ InformationTheory.klDiv (ProbabilityTheory.gaussianReal mโ vโ) (ProbabilityTheory.gaussianReal mโ vโ) โ โคKL tensorization for klDivReal over an arbitrary finite product (general Fintype index): the real-valued analogue of klDiv_pi'.
โ {ฮน : Type u_3} [inst : Fintype ฮน] {X : ฮน โ Type u_4} {mX : (i : ฮน) โ MeasurableSpace (X i)}
(P Q : (i : ฮน) โ MeasureTheory.Measure (X i)) [โ (i : ฮน), MeasureTheory.IsProbabilityMeasure (P i)]
[โ (i : ฮน), MeasureTheory.IsProbabilityMeasure (Q i)],
(โ (i : ฮน), (P i).AbsolutelyContinuous (Q i)) โ
(โ (i : ฮน), InformationTheory.klDiv (P i) (Q i) โ โค) โ
InformationTheory.klDivReal (MeasureTheory.Measure.pi P) (MeasureTheory.Measure.pi Q) =
โ i, InformationTheory.klDivReal (P i) (Q i)KL tensorization over an arbitrary finite product (general Fintype index). Reduces to the Fin-indexed klDiv_pi by reindexing through ฮน โ Fin (card ฮน).
โ {ฮน : Type u_3} [inst : Fintype ฮน] {X : ฮน โ Type u_4} {mX : (i : ฮน) โ MeasurableSpace (X i)}
(P Q : (i : ฮน) โ MeasureTheory.Measure (X i)) [โ (i : ฮน), MeasureTheory.IsProbabilityMeasure (P i)]
[โ (i : ฮน), MeasureTheory.IsProbabilityMeasure (Q i)],
InformationTheory.klDiv (MeasureTheory.Measure.pi P) (MeasureTheory.Measure.pi Q) =
โ i, InformationTheory.klDiv (P i) (Q i)KL tensorization over a finite product, indexed by Fin (n+1): klDiv (Measure.pi P) (Measure.pi Q) = โ i, klDiv (P i) (Q i).
โ {n : โ} {X : Fin n โ Type u_3} {mX : (i : Fin n) โ MeasurableSpace (X i)}
(P Q : (i : Fin n) โ MeasureTheory.Measure (X i)) [โ (i : Fin n), MeasureTheory.IsProbabilityMeasure (P i)]
[โ (i : Fin n), MeasureTheory.IsProbabilityMeasure (Q i)],
InformationTheory.klDiv (MeasureTheory.Measure.pi P) (MeasureTheory.Measure.pi Q) =
โ i, InformationTheory.klDiv (P i) (Q i)Binary KL tensorization. For probability measures, klDiv (ฮผ.prod ฮฝ) (ฮผ'.prod ฮฝ') = klDiv ฮผ ฮผ' + klDiv ฮฝ ฮฝ'.
โ {ฮฑ : Type u_1} {ฮฒ : Type u_2} {mฮฑ : MeasurableSpace ฮฑ} {mฮฒ : MeasurableSpace ฮฒ} (ฮผ ฮผ' : MeasureTheory.Measure ฮฑ)
(ฮฝ ฮฝ' : MeasureTheory.Measure ฮฒ) [MeasureTheory.IsProbabilityMeasure ฮผ] [MeasureTheory.IsProbabilityMeasure ฮผ']
[MeasureTheory.IsProbabilityMeasure ฮฝ] [MeasureTheory.IsProbabilityMeasure ฮฝ'],
InformationTheory.klDiv (ฮผ.prod ฮฝ) (ฮผ'.prod ฮฝ') = InformationTheory.klDiv ฮผ ฮผ' + InformationTheory.klDiv ฮฝ ฮฝ'Two products sharing the same first marginal reduce to the second-coordinate KL: klDiv (ฮผ.prod ฮฝ) (ฮผ.prod ฮฝ') = klDiv ฮฝ ฮฝ'.
โ {ฮฑ : Type u_1} {ฮฒ : Type u_2} {mฮฑ : MeasurableSpace ฮฑ} {mฮฒ : MeasurableSpace ฮฒ} (ฮผ : MeasureTheory.Measure ฮฑ)
(ฮฝ ฮฝ' : MeasureTheory.Measure ฮฒ) [MeasureTheory.IsProbabilityMeasure ฮผ] [MeasureTheory.IsFiniteMeasure ฮฝ]
[MeasureTheory.IsFiniteMeasure ฮฝ'], InformationTheory.klDiv (ฮผ.prod ฮฝ) (ฮผ.prod ฮฝ') = InformationTheory.klDiv ฮฝ ฮฝ'KL invariance under a measurable equivalence (โโฅ0โ-valued). For finite measures ฮผ, ฮฝ and a measurable equivalence e : ฮฑ โแต ฮฒ, klDiv (ฮผ.map e) (ฮฝ.map e) = klDiv ฮผ ฮฝ.
โ {ฮฑ : Type u_1} {ฮฒ : Type u_2} [inst : MeasurableSpace ฮฑ] [inst_1 : MeasurableSpace ฮฒ] (e : ฮฑ โแต ฮฒ)
(ฮผ ฮฝ : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsFiniteMeasure ฮผ] [MeasureTheory.IsFiniteMeasure ฮฝ],
InformationTheory.klDiv (MeasureTheory.Measure.map (โe) ฮผ) (MeasureTheory.Measure.map (โe) ฮฝ) =
InformationTheory.klDiv ฮผ ฮฝThe log-likelihood ratio is invariant (a.e.) under pushforward by a measurable equivalence: llr (ฮผ.map e) (ฮฝ.map e) (e x) =แต[ฮฝ] llr ฮผ ฮฝ x.
โ {ฮฑ : Type u_1} {ฮฒ : Type u_2} [inst : MeasurableSpace ฮฑ] [inst_1 : MeasurableSpace ฮฒ] (e : ฮฑ โแต ฮฒ)
(ฮผ ฮฝ : MeasureTheory.Measure ฮฑ) [MeasureTheory.SigmaFinite ฮผ] [MeasureTheory.SigmaFinite ฮฝ],
(fun x => MeasureTheory.llr (MeasureTheory.Measure.map (โe) ฮผ) (MeasureTheory.Measure.map (โe) ฮฝ) (e x)) =แต[ฮฝ]
MeasureTheory.llr ฮผ ฮฝFor probability measures with P โช Q, the โ-valued KL equals (Mathlib.klDiv P Q).toReal. Both measures have total mass 1, so the Mathlib correction term vanishes.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsProbabilityMeasure P]
[MeasureTheory.IsProbabilityMeasure Q],
P.AbsolutelyContinuous Q โ InformationTheory.klDivReal P Q = (InformationTheory.klDiv P Q).toRealThe integrand log ((P.rnDeriv Q x).toReal) is definitionally Mathlib's log-likelihood ratio llr P Q x.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ),
(fun x => Real.log (P.rnDeriv Q x).toReal) = MeasureTheory.llr P QClosed form of the 1-D Gaussian-Gaussian KL divergence. For variance parameters vโ โ 0, vโ โ 0, KL( N(mโ, vโ) โ N(mโ, vโ) ) = ยฝ ( log(vโ/vโ) + (vโ + (mโ-mโ)ยฒ)/vโ - 1 ).
โ {mโ mโ : โ} {vโ vโ : NNReal},
vโ โ 0 โ
vโ โ 0 โ
InformationTheory.klDivReal (ProbabilityTheory.gaussianReal mโ vโ) (ProbabilityTheory.gaussianReal mโ vโ) =
1 / 2 * (Real.log (โvโ / โvโ) + (โvโ + (mโ - mโ) ^ 2) / โvโ - 1)The Radon-Nikodym derivative of one real Gaussian w.r.t. another is, volume-a.e., the ratio of their densities.
โ {mโ mโ : โ} {vโ vโ : NNReal},
vโ โ 0 โ
vโ โ 0 โ
(ProbabilityTheory.gaussianReal mโ vโ).rnDeriv (ProbabilityTheory.gaussianReal mโ vโ) =แต[MeasureTheory.volume]
fun x => ProbabilityTheory.gaussianPDF mโ vโ x / ProbabilityTheory.gaussianPDF mโ vโ xlog (gaussianPDFReal m v x) = -ยฝ log(2ฯv) - (x - m)ยฒ/(2v) for v โ 0.
โ {mโ : โ} {vโ : NNReal},
vโ โ 0 โ
โ (x : โ),
Real.log (ProbabilityTheory.gaussianPDFReal mโ vโ x) =
-(1 / 2) * Real.log (2 * Real.pi * โvโ) - (x - mโ) ^ 2 / (2 * โvโ)โซ (x - mโ)ยฒ โgaussianReal mโ vโ = vโ + (mโ - mโ)ยฒ (bias-variance split).
โ {mโ mโ : โ} {vโ : NNReal}, โซ (x : โ), (x - mโ) ^ 2 โProbabilityTheory.gaussianReal mโ vโ = โvโ + (mโ - mโ) ^ 2โซ (x - mโ) โgaussianReal mโ vโ = 0.
โ {mโ : โ} {vโ : NNReal}, โซ (x : โ), x - mโ โProbabilityTheory.gaussianReal mโ vโ = 0โซ (x - mโ)ยฒ โgaussianReal mโ vโ = vโ.
โ {mโ : โ} {vโ : NNReal}, โซ (x : โ), (x - mโ) ^ 2 โProbabilityTheory.gaussianReal mโ vโ = โvโKL invariance under a measurable equivalence. For ฯ-finite measures P, Q on ฮฑ and a measurable equivalence e : ฮฑ โแต ฮฒ, klDivReal (P.map e) (Q.map e) = klDivReal P Q.
โ {ฮฑ : Type u_1} {ฮฒ : Type u_2} [inst : MeasurableSpace ฮฑ] [inst_1 : MeasurableSpace ฮฒ] (e : ฮฑ โแต ฮฒ)
(P Q : MeasureTheory.Measure ฮฑ) [MeasureTheory.SigmaFinite P] [MeasureTheory.SigmaFinite Q],
InformationTheory.klDivReal (MeasureTheory.Measure.map (โe) P) (MeasureTheory.Measure.map (โe) Q) =
InformationTheory.klDivReal P Qdet C(Sโ,Sโ) = det Sโ / det Sโ.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] {Sโ Sโ : Matrix ฮน ฮน โ},
Sโ.PosDef โ
Sโ.PosDef โ ((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ * ((CFC.sqrt Sโ)โปยน * CFC.sqrt Sโ).conjTranspose).det = Sโ.det / Sโ.detFor a positive-definite matrix S, the functional-calculus square root has invertible determinant: (CFC.sqrt S).det ^ 2 = S.det > 0, so the sqrt is itself invertible.
โ {ฮน : Type u_1} [inst : Fintype ฮน] [inst_1 : DecidableEq ฮน] (S : Matrix ฮน ฮน โ), S.PosDef โ IsUnit (CFC.sqrt S).detSโ.PosDef
Sโ.PosDef
โ-valued KL divergence. Returns 0 when P is not absolutely continuous with respect to Q (by convention; the โโฅ0โ-valued klDiv returns โค in that case).
{ฮฑ : Type u_1} โ [inst : MeasurableSpace ฮฑ] โ MeasureTheory.Measure ฮฑ โ MeasureTheory.Measure ฮฑ โ โThe defining affine map z โฆ m + โS z of multivariateGaussian m S, as a measurable equivalence for positive-definite S (whose functional-calculus square root is invertible).
{ฮน : Type u_1} โ
[Fintype ฮน] โ
[DecidableEq ฮน] โ EuclideanSpace โ ฮน โ {S : Matrix ฮน ฮน โ} โ S.PosDef โ EuclideanSpace โ ฮน โแต EuclideanSpace โ ฮนAn invertible matrix as a continuous linear equivalence on EuclideanSpace โ ฮน.
{ฮน : Type u_1} โ
[inst : Fintype ฮน] โ
[inst_1 : DecidableEq ฮน] โ (M : Matrix ฮน ฮน โ) โ IsUnit M.det โ EuclideanSpace โ ฮน โL[โ] EuclideanSpace โ ฮนThe Euclidean identification toLp 2 : (ฮน โ โ) โ EuclideanSpace โ ฮน as a measurable equivalence, used to transport the tensorization of Measure.pi onto EuclideanSpace.
{ฮน : Type u_1} โ [Fintype ฮน] โ (ฮน โ โ) โแต EuclideanSpace โ ฮนAn analytic non-Borel subset of โ: the image of the Baire-space witness under the continuous injection. Analyticity transfers along the continuous image; non-Borelness transfers back along the injective preimage.
โ A, MeasureTheory.AnalyticSet A โง ยฌMeasurableSet A
โ A, MeasureTheory.AnalyticSet A โง ยฌMeasurableSet A
An analytic non-Borel subset of Baire space: the projection of the diagonal witness.
โ A, MeasureTheory.AnalyticSet A โง ยฌMeasurableSet A
A closed set with a non-Borel projection. Diagonalize the universal closed set of (โ โ โ) ร (โ โ โ): were the projection Borel, its complement would be analytic, hence the projection of a closed set, hence a section of the universal set โ and evaluating that section at its own parameter is contradictory.
โ D, IsClosed D โง ยฌMeasurableSet {x | โ y, (x, y) โ D}A universal closed set. Every second-countable space X carries a closed subset of X ร (โ โ โ) whose sections run through all closed subsets of X: enumerate a countable basis together with โ
, and let the parameter select which basis elements to exclude.
โ (X : Type u_1) [inst : TopologicalSpace X] [SecondCountableTopology X],
โ S, IsClosed S โง โ (C : Set X), IsClosed C โ โ y, {x | (x, y) โ S} = CFunction.Injective MeasureTheory.embedBaireReal
Function.Injective MeasureTheory.baireMarkerBits
โ (x : โ โ โ), StrictMono (MeasureTheory.baireMarkers x)
Continuous MeasureTheory.embedBaireReal
Continuous (Cardinal.cantorFunction (1 / 3))
Continuous MeasureTheory.baireMarkerBits
The marker bits: the indicator stream of the marker set.
(โ โ โ) โ โ โ Bool
The marker sequence of x : โ โ โ: the strictly increasing sequence n + 1 + โ_{k โค n} x k, whose successive gaps encode x.
(โ โ โ) โ โ โ โ
The embedding of Baire space into โ: marker bits into the base-3 expansion.
(โ โ โ) โ โ
Pinsker's inequality with the sharp constant. tvDistReal P Q โค sqrt(klDivReal P Q / 2) for probability measures P โช Q with finite KL divergence.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ)
[inst_1 : MeasureTheory.IsProbabilityMeasure P] [inst_2 : MeasureTheory.IsProbabilityMeasure Q],
P.AbsolutelyContinuous Q โ
MeasureTheory.Integrable (MeasureTheory.llr P Q) P โ
MeasureTheory.tvDistReal P Q โค โ(InformationTheory.klDivReal P Q / 2)โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ)
[inst_1 : MeasureTheory.IsProbabilityMeasure P] [inst_2 : MeasureTheory.IsProbabilityMeasure Q],
P.AbsolutelyContinuous Q โ
MeasureTheory.Integrable (MeasureTheory.llr P Q) P โ
MeasureTheory.tvDistReal P Q โค โ(InformationTheory.klDivReal P Q / 2)โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsFiniteMeasure P]
[MeasureTheory.IsFiniteMeasure Q], {x | โ A, MeasurableSet A โง x = |(P A).toReal - (Q A).toReal|}.NonemptyPer-set squared bound. 2 (P(A) โ Q(A))ยฒ โค klDivReal P Q for any measurable set A, given P โช Q and finite KL.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsProbabilityMeasure P]
[MeasureTheory.IsProbabilityMeasure Q],
P.AbsolutelyContinuous Q โ
MeasureTheory.Integrable (MeasureTheory.llr P Q) P โ
โ (A : Set ฮฑ), MeasurableSet A โ 2 * (P.real A - Q.real A) ^ 2 โค InformationTheory.klDivReal P QData processing inequality for indicators. For probability measures P โช Q with finite KL divergence and any measurable set A, the binary KL between the marginals on (A, Aแถ) is bounded by the full KL: klBin(P(A), Q(A)) โค klDivReal P Q.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsProbabilityMeasure P]
[MeasureTheory.IsProbabilityMeasure Q],
P.AbsolutelyContinuous Q โ
MeasureTheory.Integrable (MeasureTheory.llr P Q) P โ
โ (A : Set ฮฑ), MeasurableSet A โ InformationTheory.klBin (P.real A) (Q.real A) โค InformationTheory.klDivReal P QJensen on a subset. For a finite measure Q and a measurable set A with positive mass, if f and klFun โ f are integrable and f โฅ 0 almost everywhere, then Q(A) ยท klFun((1/Q(A)) ยท โซ_A f dQ) โค โซ_A klFun(f) dQ.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] {Q : MeasureTheory.Measure ฮฑ} [MeasureTheory.IsFiniteMeasure Q] (f : ฮฑ โ โ),
MeasureTheory.Integrable f Q โ
MeasureTheory.Integrable (fun x => InformationTheory.klFun (f x)) Q โ
(โแต (x : ฮฑ) โQ, 0 โค f x) โ
โ {A : Set ฮฑ},
0 < Q.real A โ
Q.real A * InformationTheory.klFun ((Q.real A)โปยน * โซ (x : ฮฑ) in A, f x โQ) โค
โซ (x : ฮฑ) in A, InformationTheory.klFun (f x) โQAlgebraic identity. The binary KL factors through klFun: klBin p q = q ยท klFun(p/q) + (1 โ q) ยท klFun((1 โ p)/(1 โ q)).
โ (p q : โ),
0 < q โ
q < 1 โ
InformationTheory.klBin p q =
q * InformationTheory.klFun (p / q) + (1 - q) * InformationTheory.klFun ((1 - p) / (1 - q))Binary Pinsker inequality. 2 (p โ q)ยฒ โค klBin p q for p โ [0, 1], q โ (0, 1). Sharp constant 2.
โ (p q : โ), 0 โค p โ p โค 1 โ 0 < q โ q < 1 โ 2 * (p - q) ^ 2 โค InformationTheory.klBin p q
2 (1 - q)ยฒ โค -log q for q โ (0, 1]. Substitute r = 1 - q into the previous.
โ (q : โ), 0 < q โ q โค 1 โ 2 * (1 - q) ^ 2 โค -Real.log q
2 qยฒ โค -log(1 - q) for q โ [0, 1).
โ (q : โ), 0 โค q โ q < 1 โ 2 * q ^ 2 โค -Real.log (1 - q)
h(q) = -log(1-q) - 2 qยฒ is monotone on [0, 1).
MonotoneOn (fun q => -Real.log (1 - q) - 2 * q ^ 2) (Set.Ico 0 1)
Derivative of h(q) = -log(1-q) - 2 qยฒ at q < 1.
โ q < 1, HasDerivAt (fun q => -Real.log (1 - q) - 2 * q ^ 2) ((2 * q - 1) ^ 2 / (1 - q)) q
g(q) = klBin p q โ 2 (p โ q)ยฒ is monotone on [p, 1).
โ (p : โ), 0 < p โ p < 1 โ MonotoneOn (fun q => InformationTheory.klBin p q - 2 * (p - q) ^ 2) (Set.Ico p 1)
On [p, 1): derivative of g is nonnegative.
โ (p q : โ), p โค q โ 0 < q โ q < 1 โ 0 โค (q - p) * (1 - 2 * q) ^ 2 / (q * (1 - q))
Continuity of g on Ico p 1.
โ (p : โ), 0 < p โ p < 1 โ ContinuousOn (fun q => InformationTheory.klBin p q - 2 * (p - q) ^ 2) (Set.Ico p 1)
Continuity of klBin p ยท on Ico p 1.
โ (p : โ), 0 < p โ p < 1 โ ContinuousOn (fun q => InformationTheory.klBin p q) (Set.Ico p 1)
Boundary case: klBin 0 q = -log(1 - q).
โ (q : โ), InformationTheory.klBin 0 q = -Real.log (1 - q)
klBin p p = 0 for p โ (0, 1).
โ (p : โ), 0 < p โ p < 1 โ InformationTheory.klBin p p = 0
Boundary case: klBin 1 q = -log q.
โ (q : โ), InformationTheory.klBin 1 q = -Real.log q
g(q) = klBin p q โ 2 (p โ q)ยฒ is antitone on (0, p].
โ (p : โ), 0 < p โ p < 1 โ AntitoneOn (fun q => InformationTheory.klBin p q - 2 * (p - q) ^ 2) (Set.Ioc 0 p)
Factored derivative identity. Derivative of g(q) := klBin(p, q) โ 2 (p โ q)ยฒ has the factored form (q โ p) ยท (1 โ 2q)ยฒ / (q ยท (1 โ q)), reducing the sign of the derivative to sign(q โ p).
โ (p q : โ),
0 < p โ
p < 1 โ
0 < q โ
q < 1 โ
HasDerivAt (fun q => InformationTheory.klBin p q - 2 * (p - q) ^ 2)
((q - p) * (1 - 2 * q) ^ 2 / (q * (1 - q))) qDerivative of (p โ q)ยฒ with respect to q.
โ (p q : โ), HasDerivAt (fun q => (p - q) ^ 2) (-2 * (p - q)) q
On (0, p]: derivative of g is nonpositive.
โ (p q : โ), q โค p โ 0 < q โ q < 1 โ (q - p) * (1 - 2 * q) ^ 2 / (q * (1 - q)) โค 0
The derivative factor (1 โ 2q)ยฒ / (q (1 โ q)) is nonnegative on (0, 1).
โ (q : โ), 0 < q โ q < 1 โ 0 โค (1 - 2 * q) ^ 2 / (q * (1 - q))
Continuity of g on Ioc 0 p.
โ (p : โ), 0 < p โ p < 1 โ ContinuousOn (fun q => InformationTheory.klBin p q - 2 * (p - q) ^ 2) (Set.Ioc 0 p)
Continuity of fun q => (p - q)ยฒ on any set.
โ (p : โ) (s : Set โ), ContinuousOn (fun q => (p - q) ^ 2) s
Continuity of klBin p ยท on Ioc 0 p.
โ (p : โ), 0 < p โ p < 1 โ ContinuousOn (fun q => InformationTheory.klBin p q) (Set.Ioc 0 p)
Derivative of klBin p ยท at a point q โ (0, 1) with p โ (0, 1).
โ (p q : โ), 0 < p โ p < 1 โ 0 < q โ q < 1 โ HasDerivAt (fun q => InformationTheory.klBin p q) ((q - p) / (q * (1 - q))) q
Expanded form: split into constants and q-dependent pieces. Used to compute the derivative via HasDerivAt in the next shard.
โ (p q : โ),
0 < p โ
p < 1 โ
0 < q โ
q < 1 โ
InformationTheory.klBin p q =
p * Real.log p - p * Real.log q + (1 - p) * Real.log (1 - p) - (1 - p) * Real.log (1 - q)โ-valued KL is nonnegative for probability measures.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsProbabilityMeasure P]
[MeasureTheory.IsProbabilityMeasure Q], 0 โค InformationTheory.klDivReal P QFor probability measures with P โช Q, the โ-valued KL equals (Mathlib.klDiv P Q).toReal. Both measures have total mass 1, so the Mathlib correction term vanishes.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ) [MeasureTheory.IsProbabilityMeasure P]
[MeasureTheory.IsProbabilityMeasure Q],
P.AbsolutelyContinuous Q โ InformationTheory.klDivReal P Q = (InformationTheory.klDiv P Q).toRealThe integrand log ((P.rnDeriv Q x).toReal) is definitionally Mathlib's log-likelihood ratio llr P Q x.
โ {ฮฑ : Type u_1} [inst : MeasurableSpace ฮฑ] (P Q : MeasureTheory.Measure ฮฑ),
(fun x => Real.log (P.rnDeriv Q x).toReal) = MeasureTheory.llr P QP.AbsolutelyContinuous Q
MeasureTheory.Integrable (MeasureTheory.llr P Q) P
Binary KL divergence between Bernoulli(p) and Bernoulli(q).
โ โ โ โ โ
โ-valued KL divergence. Returns 0 when P is not absolutely continuous with respect to Q (by convention; the โโฅ0โ-valued klDiv returns โค in that case).
{ฮฑ : Type u_1} โ [inst : MeasurableSpace ฮฑ] โ MeasureTheory.Measure ฮฑ โ MeasureTheory.Measure ฮฑ โ โTotal variation distance between two probability measures, metric form.
Defined as the supremum of |P(A).toReal - Q(A).toReal| over measurable sets A.
{ฮฑ : Type u_1} โ
[inst : MeasurableSpace ฮฑ] โ
(P Q : MeasureTheory.Measure ฮฑ) โ [MeasureTheory.IsFiniteMeasure P] โ [MeasureTheory.IsFiniteMeasure Q] โ โd = 1 universality: every VC-1 class compresses at kernel size one, with no side information. No finiteness, no distinguished member, no chain hypothesis. Twist the class by any member to place โ
in it; there the VC bound makes co-member points comparable in the membership order, so every label set is a finite chain; anchor at its maximum, reconstruct with interConvention, and transport the scheme back with the same kernels.
โ {ฮฑ : Type u_1} [inst : DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)},
StructuralIgnorance.vcBounded ๐ 1 โ StructuralIgnorance.HasKernelScheme ๐ 1โ {ฮฑ : Type u_1} [inst : DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)},
StructuralIgnorance.vcBounded ๐ 1 โ StructuralIgnorance.HasKernelScheme ๐ 1VC bounds are twist-invariant (the direction needed for normalization).
โ {ฮฑ : Type u_1} [DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)} {d : โ},
StructuralIgnorance.vcBounded ๐ d โ
โ (Sโ : Set ฮฑ), StructuralIgnorance.vcBounded (StructuralIgnorance.twistClass ๐ Sโ) dShattering transports from the twisted class back to the original: the witnesses untwist, and the patterns correspond through the involution C โฆ C โ (B โฉ Sโ).
โ {ฮฑ : Type u_1} [DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)} {Sโ : Set ฮฑ} {B : Finset ฮฑ},
StructuralIgnorance.SetShatters (StructuralIgnorance.twistClass ๐ Sโ) B โ StructuralIgnorance.SetShatters ๐ BWith the empty set in the class, a VC bound of one makes any two points of a common member comparable in the membership order: otherwise the four assembled patterns shatter the pair.
โ {ฮฑ : Type u_1} [DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)},
StructuralIgnorance.vcBounded ๐ 1 โ
โ
โ ๐ โ
โ {a b : ฮฑ}, (โ S โ ๐, a โ S โง b โ S) โ StructuralIgnorance.memOrder ๐ a b โจ StructuralIgnorance.memOrder ๐ b aThe anchor leg of the intersection convention: a maximum of the label set in the membership order is a size-1 kernel. The upper inclusion comes from the realizing member, the lower from maximality.
โ {ฮฑ : Type u_1} [inst : DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)} {A T : Finset ฮฑ},
StructuralIgnorance.IsSample ๐ A T โ
โ {x : ฮฑ},
x โ T โ
(โ a โ T, StructuralIgnorance.memOrder ๐ a x) โ
StructuralIgnorance.IsKernel (StructuralIgnorance.interConvention ๐) 1 A T {x}Realizable label patterns are supported inside their window.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {A T : Finset ฮฑ}, StructuralIgnorance.IsSample ๐ A T โ T โ AA finite nonempty set on which the membership order is total has a maximum.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {T : Finset ฮฑ},
(โ a โ T, โ b โ T, StructuralIgnorance.memOrder ๐ a b โจ StructuralIgnorance.memOrder ๐ b a) โ
T.Nonempty โ โ x โ T, โ a โ T, StructuralIgnorance.memOrder ๐ a xโ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {a b c : ฮฑ},
StructuralIgnorance.memOrder ๐ a b โ StructuralIgnorance.memOrder ๐ b c โ StructuralIgnorance.memOrder ๐ a cโ {ฮฑ : Type u_1} (๐ : Set (Set ฮฑ)) (a : ฮฑ), StructuralIgnorance.memOrder ๐ a aTwisting by a member places the empty set in the class.
โ {ฮฑ : Type u_1} {๐ : Set (Set ฮฑ)} {Sโ : Set ฮฑ}, Sโ โ ๐ โ โ
โ StructuralIgnorance.twistClass ๐ SโKernel schemes transport back through a twist, with the same kernels.
โ {ฮฑ : Type u_1} [inst : DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)} {Sโ : Set ฮฑ} {k : โ},
StructuralIgnorance.HasKernelScheme (StructuralIgnorance.twistClass ๐ Sโ) k โ StructuralIgnorance.HasKernelScheme ๐ kโ {ฮฑ : Type u_1} (s t u : Set ฮฑ), symmDiff s t โฉ u = symmDiff (s โฉ u) (t โฉ u)โ {ฮฑ : Type u_1} [inst : DecidableEq ฮฑ] (Z T F : Finset ฮฑ), Z โฉ symmDiff T F = symmDiff (Z โฉ T) (Z โฉ F)Realizability transports through the twist, with the label set twisted inside the window.
โ {ฮฑ : Type u_1} [inst : DecidableEq ฮฑ] {๐ : Set (Set ฮฑ)} {A T : Finset ฮฑ},
StructuralIgnorance.IsSample ๐ A T โ
โ (Sโ : Set ฮฑ),
StructuralIgnorance.IsSample (StructuralIgnorance.twistClass ๐ Sโ) A
(symmDiff T (StructuralIgnorance.baseLabels Sโ A))โ {ฮฑ : Type u_1} (Sโ : Set ฮฑ) (Z : Finset ฮฑ), โ(StructuralIgnorance.baseLabels Sโ Z) = โZ โฉ SโStructuralIgnorance.vcBounded ๐ 1
A reconstruction rule: from a kept labeled pair โ kernel points and their 1-labeled part โ to a total hypothesis. Total by convention; only values on realizable pairs matter.
Type u_2 โ Type u_2
A kernel scheme of size k with no side information: one reconstruction rule under which every realizable sample contains a generating kernel of at most k points (LittlestoneโWarmuth 1986, in compression-map-free form).
{ฮฑ : Type u_1} โ [DecidableEq ฮฑ] โ Set (Set ฮฑ) โ โ โ Propฯ regenerates the sample (A, T) from the kernel Z: the kernel is kept inside the sample, has at most k points, and the reconstructed hypothesis meets A in exactly T.
{ฮฑ : Type u_1} โ [DecidableEq ฮฑ] โ StructuralIgnorance.Convention ฮฑ โ โ โ Finset ฮฑ โ Finset ฮฑ โ Finset ฮฑ โ PropThe labeled window (A, T) is realizable in ๐: some member meets A in exactly T. Points of T are labeled 1, points of A \ T are labeled 0 (LittlestoneโWarmuth 1986, realizable samples).
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ Finset ฮฑ โ Finset ฮฑ โ PropThe class ๐ shatters the finite set B: every sub-pattern of B is realized by a member. Set-grammar port of Finset.Shatters (Mathlib Combinatorics.SetFamily.Shatter).
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ Finset ฮฑ โ PropThe part of a finite window lying in the base set.
{ฮฑ : Type u_1} โ Set ฮฑ โ Finset ฮฑ โ Finset ฮฑReconstruction by intersection: the hypothesis is the intersection of all members containing the kernel's 1-labeled part. Improper โ the output need not be a member โ which is what dense chains require; the kernel's point set is not consulted beyond its labels.
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ StructuralIgnorance.Convention ฮฑThe membership order of a class: a lies below b when every member containing b contains a.
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ ฮฑ โ ฮฑ โ PropThe class relabeled by symmetric difference with a base set.
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ Set ฮฑ โ Set (Set ฮฑ)The conjugated convention: untwist the kernel's labels, reconstruct in the twisted class, twist the hypothesis back.
{ฮฑ : Type u_1} โ [DecidableEq ฮฑ] โ Set ฮฑ โ StructuralIgnorance.Convention ฮฑ โ StructuralIgnorance.Convention ฮฑVC bound in Set grammar: no shattered set exceeds d points (the port of Finset.vcDim โค d, Mathlib Combinatorics.SetFamily.Shatter).
{ฮฑ : Type u_1} โ Set (Set ฮฑ) โ โ โ PropFundamental theorem of statistical learning (5-way equivalence, BPโ ).
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool)
[MeasurableConceptClass X C],
(PACLearnable X C โ VCDim X C < โค) โง
(VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k) โง
(VCDim X C < โค โ
โ ฮต > 0,
โ mโ,
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ โ m โฅ mโ, RademacherComplexity X C D m < ฮต) โง
(PACLearnable X C โ
โ L mf,
(โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ
โ c โ C,
(MeasureTheory.Measure.pi fun x => D)
{xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal ฮต} โฅ
ENNReal.ofReal (1 - ฮด)) โง
(โ (ฮต ฮด : โ), 0 < ฮต โ 0 < ฮด โ SampleComplexity X C ฮต ฮด โค mf ฮต ฮด) โง
โ (d : โ),
VCDim X C = โd โ
โ (ฮต ฮด : โ),
0 < ฮต โ
ฮต โค 1 / 4 โ
0 < ฮด โ
ฮด โค 1 โ
ฮด โค 1 / 7 โ
1 โค d โ โ(โd - 1) / 2โโ โค SampleComplexity X C ฮต ฮด โง โ(โd - 1) / 2โโ โค mf ฮต ฮด) โง
(VCDim X C < โค โ โ d, โ (m : โ), d โค m โ GrowthFunction X C m โค โ i โ Finset.range (d + 1), m.choose i)โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool)
[MeasurableConceptClass X C],
(PACLearnable X C โ VCDim X C < โค) โง
(VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k) โง
(VCDim X C < โค โ
โ ฮต > 0,
โ mโ,
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ โ m โฅ mโ, RademacherComplexity X C D m < ฮต) โง
(PACLearnable X C โ
โ L mf,
(โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ
โ c โ C,
(MeasureTheory.Measure.pi fun x => D)
{xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal ฮต} โฅ
ENNReal.ofReal (1 - ฮด)) โง
(โ (ฮต ฮด : โ), 0 < ฮต โ 0 < ฮด โ SampleComplexity X C ฮต ฮด โค mf ฮต ฮด) โง
โ (d : โ),
VCDim X C = โd โ
โ (ฮต ฮด : โ),
0 < ฮต โ
ฮต โค 1 / 4 โ
0 < ฮด โ
ฮด โค 1 โ
ฮด โค 1 / 7 โ
1 โค d โ โ(โd - 1) / 2โโ โค SampleComplexity X C ฮต ฮด โง โ(โd - 1) / 2โโ โค mf ฮต ฮด) โง
(VCDim X C < โค โ โ d, โ (m : โ), d โค m โ GrowthFunction X C m โค โ i โ Finset.range (d + 1), m.choose i)VC characterization: C is PAC-learnable iff VCDim(C) < โ.
PROOF DECOMPOSITION: This theorem factors through the two directions above: โ : vcdim_finite_imp_uc + uc_imp_pac (in Generalization.lean) โ : pac_imp_vcdim_finite (contrapositive via double-sample)
HC at this joint: The โ direction crosses from combinatorics (VCDim, GrowthFunction) to measure theory (Measure.pi, TrueError). The โ direction crosses from measure theory back to combinatorics. Both crossings have HC > 0.
UKโ: The โ hides an ASYMMETRY: the โ proof is constructive (produces ERM), while the โ proof is non-constructive (produces hard distribution).
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool) [MeasurableConceptClass X C], PACLearnable X C โ VCDim X C < โค
Direction โ: finite VCDim implies PAC learnability.
PROOF ROUTE (via new infrastructure in Generalization.lean): Step 1: VCDim < โ โ HasUniformConvergence (vcdim_finite_imp_uc) Sub-step 1a: Sauer-Shelah gives GrowthFunction bound Sub-step 1b: Symmetrization reduces UC to growth function counting Sub-step 1c: Concentration inequality closes the bound Step 2: HasUniformConvergence โ PACLearnable (uc_imp_pac) Sub-step 2a: Construct ERM learner Sub-step 2b: ERM is consistent in realizable case Sub-step 2c: Consistent + UC โ low TrueError
KUโโ: C.Nonempty is needed for ERM but not stated as hypothesis. If C = โ , then PACLearnable is vacuously true (โ c โ C, ... is vacuous). But ERM needs a fallback hypothesis from C. Is this a genuine gap or does the empty case work out vacuously?
Counterdefinition (COUNTER-4): If the ERM approach fails for computational reasons (ERM is noncomputable, and we need a computable learner for computational learning theory), swap to the compression-based proof: VCDim < โ โ finite compression scheme (Moran-Yehudayoff 2016) โ compression scheme learner is PAC. Swap condition: When proving COMPUTATIONAL PAC learnability (polynomial time).
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool), VCDim X C < โค โ โ [MeasurableConceptClass X C], PACLearnable X C
Quantitative sample-complexity sandwich attached to any PAC witness. Packages: (1) PAC guarantee, (2) SampleComplexity โค mf, (3) NFL/VC lower bound on both SampleComplexity and mf.
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool)
[MeasurableConceptClass X C],
PACLearnable X C โ
โ L mf,
(โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ
โ c โ C,
(MeasureTheory.Measure.pi fun x => D)
{xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal ฮต} โฅ
ENNReal.ofReal (1 - ฮด)) โง
(โ (ฮต ฮด : โ), 0 < ฮต โ 0 < ฮด โ SampleComplexity X C ฮต ฮด โค mf ฮต ฮด) โง
โ (d : โ),
VCDim X C = โd โ
โ (ฮต ฮด : โ),
0 < ฮต โ
ฮต โค 1 / 4 โ
0 < ฮด โ
ฮด โค 1 โ ฮด โค 1 / 7 โ 1 โค d โ โ(โd - 1) / 2โโ โค SampleComplexity X C ฮต ฮด โง โ(โd - 1) / 2โโ โค mf ฮต ฮดAny PAC witness (L, mf) gives an upper bound on SampleComplexity: the infimum is at most the witness sample size.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) (L : BatchLearner X Bool) (mf : โ โ โ โ โ),
(โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ
โ c โ C,
(MeasureTheory.Measure.pi fun x => D)
{xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal ฮต} โฅ
ENNReal.ofReal (1 - ฮด)) โ
โ (ฮต ฮด : โ), 0 < ฮต โ 0 < ฮด โ SampleComplexity X C ฮต ฮด โค mf ฮต ฮดSample complexity lower bound: โ(d-1)/2โ โค SampleComplexity.
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool) (d : โ),
VCDim X C = โd โ
โ (ฮต ฮด : โ),
0 < ฮต โ
ฮต โค 1 / 4 โ
0 < ฮด โ
ฮด โค 1 โ
ฮด โค 1 / 7 โ
1 โค d โ
(โ h โ C, Measurable h) โ
(โ (c : Concept X Bool), Measurable c) โ
WellBehavedVC X C โ โ(โd - 1) / 2โโ โค SampleComplexity X C ฮต ฮดPAC lower bound membership: if m achieves PAC for C with VCDim = d, then m โฅ โ(d-1)/(64ฮต)โ. This is the core adversarial counting argument factored for PAC.lean assembly. Note: the tight constant is (d-1)/(2ฮต) (EHKV 1989); see EHKV.lean.
Proof route (double-averaging on shattered set): 1. VCDim = d โ โ shattered S with |S| = d 2. D = uniform on S (probability measure, each point has weight 1/d) 3. m < โ(d-1)/(64ฮต)โ โ 2m < d โ NFL counting applies 4. Double-averaging over 2^d labelings: E_f[E_xs[error]] โฅ (d-m)/(2d) > 1/4 5. Reversed Markov: โ cโ โ C with Pr[error โค 1/8] โค 6/7 6. For ฮต โค 1/8: Pr[error โค ฮต] โค 6/7 = 1 - 1/7, contradicting PAC
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool) (d : โ),
VCDim X C = โd โ
โ (ฮต ฮด : โ),
0 < ฮต โ
ฮต โค 1 / 4 โ
0 < ฮด โ
ฮด โค 1 โ
ฮด โค 1 / 7 โ
1 โค d โ
โ
m โ
{m |
โ L,
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ
โ c โ C,
(MeasureTheory.Measure.pi fun x => D)
{xs | D {x | L.learn (fun i => (xs i, c (xs i))) x โ c x} โค ENNReal.ofReal ฮต} โฅ
ENNReal.ofReal (1 - ฮด)},
โ(โd - 1) / 2โโ โค mGrowth function polynomially bounded โ VCDim < โค. Reverse direction: if GrowthFunction m โค โ_{iโคd} C(m,i) for all m โฅ d, then VCDim โค d (otherwise GrowthFunction = 2^m for m = VCDim > d).
โ (X : Type u) (C : ConceptClass X Bool), (โ d, โ (m : โ), d โค m โ GrowthFunction X C m โค โ i โ Finset.range (d + 1), m.choose i) โ VCDim X C < โค
Fundamental theorem: finite VC dim โ finite compression scheme with side information. Moran-Yehudayoff 2016 (arXiv:1503.06960). Sorry-free via Compression.lean. ฮโโ RESOLVED: CompressionSchemeWithInfo parameterized by concept class C. The no-side-info version (Littlestone-Warmuth conjecture) remains open.
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k
The forward direction of the Moran-Yehudayoff theorem: finite VC dimension implies existence of a compression scheme with finite side information.
The construction: 1. Build a proper finite-support learner L from VC + Sauer-Shelah 2. For sample S: extract c, Y = pointSupport S, HY = hypothesis envelope 3. Apply approximate minimax on the agreement game โ distribution p on HY 4. Apply VC ฮต-approximation on agreement tests โ T representative hypotheses 5. Kernel = union of witness subsets for T hypotheses 6. Side info = incidence: which hypothesis's witness contains each kernel point 7. Reconstruct by majority vote over T hypotheses
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ k cs, CompressionSchemeWithInfo.size cs = k
Finite VC dimension implies existence of a proper finite-support learner. The construction uses ERM + finite_support_vc_approx on the disagreement family.
โ (X : Type u) (C : ConceptClass X Bool), Set.Nonempty C โ VCDim X C < โค โ โ _L, True
supportError expressed in terms of boolTestExpectation of a disagreement test.
โ {X : Type u} (Y : Finset X) (q : FinitePMF โฅY) (h c : X โ Bool),
supportError Y q h c = boolTestExpectation q fun y => decide (h โy โ c โy)VC dimension of the disagreement family is bounded by VCDim(C). Restriction to Y and xor with c do not increase shattering dimension.
โ {X : Type u} [inst : DecidableEq X] (C : ConceptClass X Bool) (c : X โ Bool) (Y : Finset X) {d : โ},
VCDim X C โค โd โ (disagreementFamilyโ C c Y).boolVCDim โค dThe Moran-Yehudayoff forward construction. Uses finalizeIncidenceScheme to package the majority-vote scheme with universe-correct Info type.
The agent must provide: compressCore, blockHyp, rowHyp, hsmall, hsub, hagree, hmajor. These are the MY wiring.
โ (X : Type u) (C : ConceptClass X Bool),
Set.Nonempty C โ
โ (L : ProperFiniteSupportLearner X C), VCDim X C < โค โ โ (_K : โ), โ k cs, CompressionSchemeWithInfo.size cs = kGeneric roundtrip theorem for the hround sorry.
If:
encodeWitnessInfo is used in compressCore,decodeWitnessXCoords and decodeWitnessLabel are used in blockHyp, andthen the decoded block hypothesis is exactly the representative hypothesis.
โ {X : Type u} [inst : DecidableEq X] (learn : {m : โ} โ (Fin m โ X ร Bool) โ X โ Bool) (kernel : Finset (X ร Bool))
(c : X โ Bool) (K : โ) (W : Finset X) (h : X โ Bool),
kernel.card โค K โ
(โ x โ W, (x, c x) โ kernel) โ
(โ p โ kernel, p.2 = c p.1) โ
learn (labeledSampleOfFinset c W) = h โ
โ (x : X),
have info := encodeWitnessInfo kernel c K W;
have blockXCoords := decodeWitnessXCoords kernel info;
have blockLabel := decodeWitnessLabel kernel;
learn (labeledSampleOfFinset blockLabel blockXCoords) x = h xIf two label functions agree on all points of Z, then the labeled samples they induce on Z.equivFin are equal.
โ {X : Type u} [DecidableEq X] {โโ โโ : X โ Bool} {Z : Finset X},
(โ x โ Z, โโ x = โโ x) โ labeledSampleOfFinset โโ Z = labeledSampleOfFinset โโ ZIf every (x, c x) with x โ W lies in kernel, and kernel.card โค K, then decoding the encoded witness positions gives back exactly W.
โ {X : Type u} [inst : DecidableEq X] (kernel : Finset (X ร Bool)) (c : X โ Bool) {K : โ} (W : Finset X),
kernel.card โค K โ (โ x โ W, (x, c x) โ kernel) โ decodeWitnessXCoords kernel (encodeWitnessInfo kernel c K W) = WOn the encoded witness support, the decoded label function agrees with the true label function c, provided every pair in the kernel has the correct second coordinate.
โ {X : Type u} [inst : DecidableEq X] (kernel : Finset (X ร Bool)) (c : X โ Bool) (W : Finset X),
(โ x โ W, (x, c x) โ kernel) โ (โ p โ kernel, p.2 = c p.1) โ โ x โ W, decodeWitnessLabel kernel x = c xGenuine approximate minimax via MWU regret extraction. If every column mixture admits a pure row with expected payoff โฅ v, then there is a row mixture with payoff โฅ v - ฮต against every column.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [Nonempty R] [Nonempty C] [DecidableEq R]
[DecidableEq C] (M : R โ C โ Bool) (v ฮต : โ),
0 < ฮต โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ
โ p, โ (c : C), v - ฮต โค boolGamePayoff M p cA single weight is bounded by the potential.
โ {C : Type u_1} [inst : Fintype C] (cfg : MWUConfig C) (c : C), cfg.weights c โค cfg.potentialExact individual-weight tracking: the weight of column c after T rounds is (1-ฮท) to the number of rounds in which c was hit.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] (M : R โ C โ Bool) (ฮท : โ)
(hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) (T : โ)
(c : C), (mwuConfig M ฮท hฮท1 v hrow T).weights c = (1 - ฮท) ^ mwuHitCountโ M ฮท hฮท1 v hrow T cโ {C : Type u_1} [inst : Fintype C] (weights weights_1 : C โ โ) (e_weights : weights = weights_1)
(weights_pos : โ (c : C), 0 < weights c),
{ weights := weights, weights_pos := weights_pos } = { weights := weights_1, weights_pos := โฏ }Potential bound after T steps: ฮฆ_T โค |C| ยท (1 - ฮทv)^T.
This is the core MWU guarantee. Combined with individual weight lower bounds (w_T(c) = (1-ฮท)^{losses(c)}), it yields the regret bound.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] (M : R โ C โ Bool)
(ฮท : โ),
0 โค ฮท โ
โ (hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0)
(T : โ), (mwuConfig M ฮท hฮท1 v hrow T).potential โค โ(Fintype.card C) * (1 - ฮท * v) ^ TPotential bound after one step: ฮฆ' โค ฮฆ ยท (1 - ฮทยทv).
โ {R : Type u_1} {C : Type u_2} [Fintype R] [inst : Fintype C] [inst_1 : Nonempty C] (M : R โ C โ Bool) (ฮท : โ),
0 โค ฮท โ
โ (hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0)
(cfg : MWUConfig C), (mwuUpdateWeights M ฮท hฮท1 cfg โฏ.choose).potential โค cfg.potential * (1 - ฮท * v)Best response payoff โฅ v ยท ฮฆ in terms of weights.
โ {R : Type u_1} {C : Type u_2} [Fintype R] [inst : Fintype C] [inst_1 : Nonempty C] (M : R โ C โ Bool) (v : โ)
(hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) (cfg : MWUConfig C),
v * cfg.potential โค โ c, cfg.weights c * if M โฏ.choose c = true then 1 else 0โ {C : Type u_1} [inst : Fintype C] [Nonempty C] (cfg : MWUConfig C), 0 < cfg.potentialโ {C : Type u_1} [inst : Fintype C] (self : MWUConfig C) (c : C), 0 < self.weights cโ {C : Type u_1} [inst : Fintype C] {R : Type u_2} (M M_1 : R โ C โ Bool),
M = M_1 โ
โ (ฮท ฮท_1 : โ) (e_ฮท : ฮท = ฮท_1) (hฮท1 : ฮท < 1) (cfg cfg_1 : MWUConfig C),
cfg = cfg_1 โ โ (r r_1 : R), r = r_1 โ mwuUpdateWeights M ฮท hฮท1 cfg r = mwuUpdateWeights M_1 ฮท_1 โฏ cfg_1 r_1โ (C : Type u_1) [inst : Fintype C], (mwuInit C).potential = โ(Fintype.card C)
The minimax value of a Boolean game is at most 1.
โ {R : Type u_1} {C : Type u_2} [Fintype R] [inst : Fintype C] [Nonempty C] (M : R โ C โ Bool) (v : โ),
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ v โค 1Arithmetic core: from the potential bound and sufficiently small ฮท / large T, deduce a per-column hit-rate lower bound. Uses Real.log โ exactly 4 Mathlib lemmas.
โ {N H T : โ} {ฮท v ฮต : โ},
0 < โN โ
0 < ฮท โ
ฮท < 1 โ
v โค 1 โ
0 < T โ (1 - ฮท) ^ H โค โN * (1 - ฮท * v) ^ T โ ฮท โค ฮต / 4 โ Real.log โN / (ฮท * โT) โค ฮต / 4 โ v - ฮต โค โH / โTโ {R : Type u_1} {C : Type u_2} [inst : Fintype R] (M : R โ C โ Bool) (p : FinitePMF R) (c : C),
0 โค boolGamePayoff M p cEmpirical payoff of the MWU row sequence equals the normalized hit count.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] [inst_3 : DecidableEq R]
(M : R โ C โ Bool) (ฮท : โ) (hฮท1 : ฮท < 1) (v : โ)
(hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) {T : โ} (hT : 0 < T) (c : C),
boolGamePayoff M (empiricalPMF hT (mwuRows M ฮท hฮท1 v hrow T)) c = โ(mwuHitCountโ M ฮท hฮท1 v hrow T c) / โTThe recursive hit counter agrees with the sum of Boolean indicators over the emitted row sequence.
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : Fintype C] [inst_2 : Nonempty C] (M : R โ C โ Bool) (ฮท : โ)
(hฮท1 : ฮท < 1) (v : โ) (hrow : โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) (T : โ)
(c : C), โ(mwuHitCountโ M ฮท hฮท1 v hrow T c) = โ t, if M (mwuRows M ฮท hฮท1 v hrow T t) c = true then 1 else 0Specialized empirical-payoff identity for ApproxMinimax (avoids cyclic import with FiniteVCApprox).
โ {R : Type u_1} {C : Type u_2} [inst : Fintype R] [inst_1 : DecidableEq R] {T : โ} (hT : 0 < T) (rs : Fin T โ R)
(M : R โ C โ Bool) (c : C), boolGamePayoff M (empiricalPMF hT rs) c = (โ t, if M (rs t) c = true then 1 else 0) / โTEvery hypothesis in the envelope is in C.
โ {X : Type u} {C : ConceptClass X Bool} (L : ProperFiniteSupportLearner X C) (c : X โ Bool) (Y : Finset X),
โ h โ hypothesisEnvelope L c Y, h โ Cโ {X : Type u} {C : ConceptClass X Bool} (self : ProperFiniteSupportLearner X C) {m : โ} (S : Fin m โ X ร Bool),
self.learn S โ CFor each C-realizable sample, the proper learner provides a row-response for the minimax game on the hypothesis envelope.
โ {X : Type u} {C : ConceptClass X Bool} (L : ProperFiniteSupportLearner X C),
โ c โ C,
โ (Y : Finset X) [Nonempty โฅY] (HY : Finset (X โ Bool)),
HY = hypothesisEnvelope L c Y โ
โ (q : FinitePMF โฅY), โ h, 2 / 3 โค โ y, q.prob y * if decide (โh โy = c โy) = true then 1 else 0Weighted agreement = 1 - supportError.
โ {X : Type u} (Y : Finset X) (q : FinitePMF โฅY) (h c : X โ Bool),
(โ y, q.prob y * if h โy = c โy then 1 else 0) = 1 - supportError Y q h cโ {X : Type u} {C : ConceptClass X Bool} (self : ProperFiniteSupportLearner X C),
โ c โ C,
โ (Y : Finset X) (q : FinitePMF โฅY),
โ Z โ Y, Z.card โค self.sampleBound โง supportError Y q (self.learn (labeledSampleOfFinset c Z)) c โค 1 / 3Finite-support distributions uniformly approximate any distribution on a VC class. For a class of VC dimension at most d and any ฮต > 0, there exists T = T(d, ฮต) such that every finitely supported distribution ฮผ is within ฮต (uniformly over the class) of some empirical distribution on T points. A density-style reduction that lets the approximate minimax / MWU machinery, which lives in finite support, apply to general distributions.
โ (d : โ) (ฮต : โ),
0 < ฮต โ
โ T,
โ (hT : 0 < T),
โ {H : Type u_1} [inst : Fintype H] [inst_1 : DecidableEq H] (A : Finset (H โ Bool)),
A.boolVCDim โค d โ
โ (ฮผ : FinitePMF H),
โ hs, โ a โ A, |boolTestExpectation ฮผ a - boolTestExpectation (empiricalPMF hT hs) a| โค ฮตโ {H : Type u_1} [inst : Fintype H] [DecidableEq H] [inst_2 : MeasurableSpace H] [MeasurableSingletonClass H]
(ฮผ : FinitePMF H) (a : H โ Bool),
TrueErrorReal (H โ โ) (extendBoolโ a) (fun x => false) (PMF.map Sum.inl (FinitePMF.toPMFโ ฮผ)).toMeasure =
boolTestExpectation ฮผ aA convex combination of values in {0, 1} is nonnegative.
โ {H : Type u_1} [inst : Fintype H] (ฮผ : FinitePMF H) (f : H โ Bool), 0 โค boolTestExpectation ฮผ fA convex combination of values in {0, 1} is at most 1.
โ {H : Type u_1} [inst : Fintype H] (ฮผ : FinitePMF H) (f : H โ Bool), boolTestExpectation ฮผ f โค 1โ {H : Type u_1} [inst : Fintype H] (self : FinitePMF H) (h : H), 0 โค self.prob hโ {H : Type u_1} [inst : Fintype H] (self : FinitePMF H), โ h, self.prob h = 1Final existential wrapper: closes the theorem in the exact form expected.
โ {X : Type u} {C : ConceptClass X Bool} (T K : โ)
(compressCore : {m : โ} โ (Fin m โ X ร Bool) โ Finset (X ร Bool) ร IncidenceInfo T K)
(blockHyp : Finset (X ร Bool) โ IncidenceInfo T K โ Fin T โ X โ Bool)
(rowHyp : {m : โ} โ (S : Fin m โ X ร Bool) โ (โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) โ Fin T โ X โ Bool),
0 < T โ
(โ {m : โ} (S : Fin m โ X ร Bool), (compressCore S).1.card โค K) โ
(โ {m : โ} (S : Fin m โ X ร Bool), โ(compressCore S).1 โ Set.range S) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m) (t : Fin T),
blockHyp (compressCore S).1 (compressCore S).2 t (S i).1 = rowHyp S hreal t (S i).1) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m),
(โ t, if rowHyp S hreal t (S i).1 = (S i).2 then 1 else 0) / โT > 1 / 2) โ
โ k cs, CompressionSchemeWithInfo.size cs = kโ {X : Type u} {inst : DecidableEq X} [inst_1 : DecidableEq X] (kernel kernel_1 : Finset (X ร Bool)),
kernel = kernel_1 โ
โ (c c_1 : X โ Bool),
c = c_1 โ
โ (K : โ) (W W_1 : Finset X), W = W_1 โ encodeWitnessInfo kernel c K W = encodeWitnessInfo kernel_1 c_1 K W_1Bridges the FinitePMF view and the sample-average view: the expectation of a Bool-valued test under the empirical PMF of a sample equals the sample average (1/T) โ_t f (s_t). This lets the MWU updates and the approximation transfer principle live in the same distributional framework.
โ {H : Type u_1} [inst : Fintype H] [inst_1 : DecidableEq H] {T : โ} (hT : 0 < T) (hs : Fin T โ H) (f : H โ Bool),
boolTestExpectation (empiricalPMF hT hs) f = (โ t, if f (hs t) = true then 1 else 0) / โTIdentifies the game-theoretic payoff (a row distribution against a fixed column in the Bool game) with the corresponding test expectation. The translation that lets the MWU regret bound be applied directly to the compression problem.
โ {R : Type u_1} [inst : Fintype R] [DecidableEq R] {C : Type u_2} (M : R โ C โ Bool) (p : FinitePMF R) (c : C),
boolGamePayoff M p c = boolTestExpectation p fun r => M r cVC dimension of the agreement-test family is bounded by 2^(d+1) - 1, where d bounds the VC dimension of the concept class C. Uses Assouad's coding argument directly: if a shattered set T in โฅHY has |T| โฅ 2^(d+1), embed bitstrings into T, extract d+1 distinct points from Y via shattering, and show these points are shattered by C (using the XOR trick where b(j) = decide(g(x_j) = c(x_j)) absorbs the agree/disagree flip).
โ {X : Type u} [DecidableEq X] (C : ConceptClass X Bool) (c : X โ Bool) (Y : Finset X) (HY : Finset (X โ Bool)),
(โ h โ HY, h โ C) โ โ {d : โ}, VCDim X C โค โd โ (agreeTests c Y HY).boolVCDim โค 2 ^ (d + 1) - 1Compression with side info implies finite VC dimension. Proof by pigeonhole: compress is injective on C-realizable labelings (by correctness), but compressed outputs form a bounded set.
โ (X : Type u) (C : ConceptClass X Bool), (โ k cs, cs.size = k) โ VCDim X C < โค
โ {X : Type u} {C : ConceptClass X Bool} {S T : Finset X}, T โ S โ Shatters X C S โ Shatters X C TExponential beats polynomial for the compression pigeonhole argument.
โ (s : โ), (s + 1) ^ 2 * (4 * (s + 1) ^ 2) ^ s < 2 ^ (2 * (s + 1) * (s + 1))
โ (k : โ), k + 1 โค 2 ^ k
Pigeonhole core: if two C-realizable samples over the same points with different labelings produce the same (kernel, info) pair, correctness forces the labelings to agree.
โ {X : Type u} {n : โ} {C : ConceptClass X Bool} (cs : CompressionSchemeWithInfo X Bool C) (pts : Fin n โ X),
Function.Injective pts โ
โ (f g : Fin n โ Bool),
(โ c โ C, โ (i : Fin n), c (pts i) = f i) โ
(โ c โ C, โ (i : Fin n), c (pts i) = g i) โ
((cs.compress fun i => (pts i, f i)) = cs.compress fun i => (pts i, g i)) โ f = gCorrectness: reconstructed hypothesis agrees with every sample point, when the sample is C-realizable
โ {X : Type u} {Y : Type v} {C : ConceptClass X Y} (self : CompressionSchemeWithInfo X Y C) {m : โ} (S : Fin m โ X ร Y),
(โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) โ
โ (i : Fin m), self.reconstruct (self.compress S).1 (self.compress S).2 (S i).1 = (S i).2Compressed set is a subset of the sample
โ {X : Type u} {Y : Type v} {C : ConceptClass X Y} (self : CompressionSchemeWithInfo X Y C) {m : โ} (S : Fin m โ X ร Y),
โ(self.compress S).1 โ Set.range SCompressed set is small
โ {X : Type u} {Y : Type v} {C : ConceptClass X Y} (self : CompressionSchemeWithInfo X Y C) {m : โ} (S : Fin m โ X ร Y),
(self.compress S).1.card โค self.kernelSizeFundamental theorem: Rademacher complexity characterization. BPโ : two asymmetric directions crossing different paradigm joints. Uses uniform vanishing (โ mโ โ D), which is the textbook-standard form.
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool)
[MeasurableConceptClass X C],
PACLearnable X C โ
โ ฮต > 0,
โ mโ,
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ โ m โฅ mโ, RademacherComplexity X C D m < ฮตVCDim finite โ Rademacher vanishes uniformly. The bound mโ depends only on d and ฮต, NOT on D.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool),
VCDim X C < โค โ
โ ฮต > 0,
โ mโ,
โ (D : MeasureTheory.Measure X),
MeasureTheory.IsProbabilityMeasure D โ โ m โฅ mโ, RademacherComplexity X C D m < ฮตโ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) (D : MeasureTheory.Measure X),
VCDim X C = โ0 โ
โ (m : โ),
0 < m โ
โ [MeasureTheory.IsProbabilityMeasure (MeasureTheory.Measure.pi fun x => D)],
RademacherComplexity X C D m โค 1 / โโmWhen VCDim = 0, Rademacher complexity is bounded by 1/โm.
VCDim = 0 means no singleton is shattered, so the concept class acts as a single effective labeling. EmpRad โค 1/โm by Khintchine's inequality / Jensen. This avoids the d > 0 hypothesis of vcdim_bounds_rademacher_quantitative.
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C = โ0 โ โ (hโ hโ : Concept X Bool), hโ โ C โ hโ โ C โ โ (x : X), hโ x = hโ x
VC dimension upper bounds Rademacher complexity: Rad โค โ(2dยทlog(em/d)/m).
The proof decomposes into: (1) Pointwise: EmpRad(xs) โค B for all xs [Massart + Sauer-Shelah] (2) Integral: Rad = โซ EmpRad โค โซ B = B [probability measure]
Step (2) is proved. Step (1) for B โฅ 1 follows from EmpRad โค 1. Step (1) for B < 1 requires Massart finite lemma + Sauer-Shelah growth bound.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) (D : MeasureTheory.Measure X) (m : โ),
0 < m โ
โ (d : โ),
VCDim X C = โd โ
0 < d โ
d โค m โ
โ [MeasureTheory.IsProbabilityMeasure (MeasureTheory.Measure.pi fun x => D)],
RademacherComplexity X C D m โค โ(2 * โd * Real.log (Real.exp 1 * โm / โd) / โm)For a Set-based concept class C with VCDim X C = d, the number of distinct restrictions of C to any finite set S is bounded by โ_{iโคd} C(|S|, i).
This bridges from our Set-based VCDim to Mathlib's Finset.vcDim on the restriction to S, using the fact that โฅS is Fintype for any Finset S.
โ {X : Type u} (C : ConceptClass X Bool) (S : Finset X) (d : โ),
VCDim X C = โd โ {f | โ c โ C, โ (x : โฅS), c โx = f x}.ncard โค โ i โ Finset.range (d + 1), S.card.choose iMassart finite lemma: E_ฯ[max_{j โค N} Z_j] โค ฯโ(2 log N).
โ {m : โ},
0 < m โ
โ {N : โ} (hN : 0 < N) (Z : Fin N โ SignVector m โ โ) (ฯ_param : โ),
0 < ฯ_param โ
(โ (j : Fin N) (t : โ),
0 โค t โ
1 / โ(Fintype.card (SignVector m)) * โ sv, Real.exp (t * Z j sv) โค Real.exp (t ^ 2 * ฯ_param ^ 2 / 2)) โ
(1 / โ(Fintype.card (SignVector m)) * โ sv, Finset.univ.sup' โฏ fun j => Z j sv) โค ฯ_param * โ(2 * Real.log โN)Soft-max bound: exp(t ยท Finset.sup') โค ฮฃ exp(t ยท f_i).
โ {ฮน : Type u_1} [DecidableEq ฮน] (s : Finset ฮน) (hs : s.Nonempty) (f : ฮน โ โ) (t : โ),
0 โค t โ Real.exp (t * s.sup' hs f) โค โ i โ s, Real.exp (t * f i)โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) (D : MeasureTheory.Measure X) (m : โ), 0 < m โ โ [MeasureTheory.IsProbabilityMeasure (MeasureTheory.Measure.pi fun x => D)], RademacherComplexity X C D m โค 1
Analytical lemma: for d > 0, m โฅ โ32(d+1)/ฮตโดโ+1, ฮต โ (0,1], we have 2dยทlog(em/d)/m < ฮตยฒ.
Uses Real.log_le_rpow_div with exponent 1/2: log(x) โค x^(1/2)/(1/2) = 2โx. Then 2dยทlog(em/d)/m โค 2dยท2โ(em/d)/(m) โค ฮตยฒ.
โ (d m : โ) (ฮต : โ),
0 < ฮต โ
ฮต โค 1 โ 0 < d โ d โค m โ โ32 * (โd + 1) / ฮต ^ 4โโ + 1 โค m โ 2 * โd * Real.log (Real.exp 1 * โm / โd) / โm < ฮต ^ 2VCDim < โค โ PACLearnable via UC route.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool),
VCDim X C < โค โ
(โ h โ C, Measurable h) โ (โ (c : Concept X Bool), Measurable c) โ WellBehavedVC X C โ PACLearnable X CFinite VCDim implies uniform convergence. Proof: VCDim < โ โ UC.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool),
VCDim X C < โค โ
(โ h โ C, Measurable h) โ (โ (c : Concept X Bool), Measurable c) โ WellBehavedVC X C โ HasUniformConvergence X CVCDim < โค โ growth function polynomially bounded by partial binomial sum. Forward direction of fundamental_theorem conjunct 5. Uses Sauer-Shelah: GrowthFunction(m) โค โ_{iโคd} C(m,i) where d = VCDim.
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ d, โ (m : โ), d โค m โ GrowthFunction X C m โค โ i โ Finset.range (d + 1), m.choose i
UC bad-event bound: for m โฅ mโ(v,ฮต,ฮด), the probability of the bad event (โ h with |TrueErr-EmpErr| โฅ ฮต) is at most ฮด. Composes symmetrization_uc_bound with growth_exp_le_delta.
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
ฮด < 1 โ
โ (v : โ),
0 < v โ
(โ (n : โ), v โค n โ GrowthFunction X C n โค โ i โ Finset.range (v + 1), n.choose i) โ
(16 * Real.exp 1 * (โv + 1) / ฮต ^ 2) ^ (v + 1) / ฮด โค โm โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
|TrueErrorReal X h c D -
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮต} โค
ENNReal.ofReal ฮดThe symmetrization uniform convergence bound: two-sided version. P[โhโC: |TrueErr-EmpErr| โฅ ฮต] โค 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8).
Proof strategy (4 steps):
1. Decompose absolute value: |TrueErr - EmpErr| โฅ ฮต โ (TrueErr - EmpErr โฅ ฮต) โจ (EmpErr - TrueErr โฅ ฮต)
have abs_decomp : โ (a b : โ),
|a - b| โฅ ฮต โ a - b โฅ ฮต โจ b - a โฅ ฮต := by
intro a b; constructor
ยท intro h; by_cases h' : a - b โฅ ฮต
ยท exact Or.inl h'
ยท exact Or.inr (by linarith [abs_sub_comm a b, le_abs_self (a - b)])
ยท intro h; cases h with
| inl h => exact le_trans (le_of_eq (abs_of_nonneg (by linarith))) (by linarith)
| inr h => exact le_trans (le_of_eq (abs_of_nonpos (by linarith) โธ ...)) ...2. Upper tail: P[โhโC: TrueErr-EmpErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step + double_sample_pattern_bound.3. Lower tail: P[โhโC: EmpErr-TrueErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step to the event EmpErr-TrueErr โฅ ฮต and bound the double-sample event {EmpErr_S - EmpErr_{S'} โฅ ฮต/2}. have swap_symmetry : DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S) - EmpErr(S') โฅ ฮต/2} = DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S') - EmpErr(S) โฅ ฮต/2} := Measure.prod_swap ... 4. Union bound: P[|gap| โฅ ฮต] โค P[gap โฅ ฮต] + P[gap โค -ฮต] โค 2ยทGFยทexp(...) + 2ยทGFยทexp(...) = 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
-- Uses: MeasureTheory.measure_union_le for the union of two events
-- CAST: 2 * X + 2 * X = 4 * X in ENNReal (need ENNReal.add_mul or similar)References: SSBD Theorem 6.7, Kakade-Tewari Lecture 19
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
|TrueErrorReal X h c D -
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮต} โค
ENNReal.ofReal (4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))Symmetrization step for the lower tail: P[โh: EmpErr-TrueErr โฅ ฮต] โค 2ยทP_{double}[โh: EmpErr_S-EmpErr_{S'} โฅ ฮต/2].
Mirror of symmetrization_step for the opposite direction. Uses hoeffding_one_sided_upper instead of hoeffding_one_sided.
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) - TrueErrorReal X h c D โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}Symmetrization: the probability of a large gap TrueErr-EmpErr is at most twice the probability of a large gap EmpErr'-EmpErr on the double sample.
Proof strategy (6 steps):
1. Witness selection: For S in the bad event, โh* โ C with TrueErr(h) - EmpErr_S(h) โฅ ฮต.
-- In the bad event set, extract h* by classical choice
have h_witness : โ xs โ bad_event, โ h* โ C,
TrueErrorReal X h* c D - EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โฅ ฮต2. Ghost sample mean: E_{S'}[EmpErr_{S'}(h)] = TrueErr(h) โฅ EmpErr_S(h*) + ฮต.
MeasureTheory.integral_pi to compute E[EmpErr] over product measure. have expected_emp_err : โ h* : Concept X Bool, โซ xs, EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โ(Measure.pi (fun _ : Fin m => D)) = TrueErrorReal X h* c D := by ... 3. Hoeffding on ghost sample: P_{S'}[EmpErr_{S'}(h) < TrueErr(h) - ฮต/2] โค exp(-mฮตยฒ/2).
hoeffding_one_sided with t = ฮต/2.hm_large hypothesis ensures exp(-mฮตยฒ/2) < 1/2: 2ยทln2 โค mฮตยฒ โน mฮตยฒ/2 โฅ ln2 โน exp(-mฮตยฒ/2) โค 1/2. have hoeffding_ghost : โ h* โ C, Measure.pi (fun _ : Fin m => D) {xs' | EmpiricalError X Bool h* (sample xs') (zeroOneLoss Bool) < TrueErrorReal X h* c D - ฮต/2} โค ENNReal.ofReal (Real.exp (-m * (ฮต/2)^2 * 2)) := by intro h* _; exact hoeffding_one_sided D h* c m hm (ฮต/2) (by linarith) (by ...) (by ...) 4. Complementary probability: P_{S'}[EmpErr_{S'}(h) - EmpErr_S(h) โฅ ฮต/2] โฅ 1/2.
5. Conditional to unconditional: The witness h* from step 1 also witnesses the double-sample event โhโC: EmpErr'-EmpErr โฅ ฮต/2. So: P_{S'}[double event | S bad] โฅ 1/2.
have conditional_bound : โ xs โ bad_event,
Measure.pi (fun _ : Fin m => D)
{xs' | โ h โ C, EmpiricalError ... xs' - EmpiricalError ... xs โฅ ฮต/2}
โฅ ENNReal.ofReal (1/2) := by ...6. Fubini integration: By Measure.prod_apply and Fubini: P_{S,S'}[double event] = โซ_S P_{S'}[double event | S] โฅ (1/2) ยท P_S[bad event] โน P_S[bad event] โค 2 ยท P_{S,S'}[double event].
-- Uses: MeasureTheory.Measure.prod_apply or lintegral_prod
-- MEASURABILITY: the double-sample event is measurable as a finite union
-- of sets of the form {(xs,xs') | EmpErr'(h) - EmpErr(h) โฅ ฮต/2} for h โ C.
-- Since C may be infinite, measurability requires care: the sup over h
-- must be shown to be measurable. For finite restriction patterns (โค 2^m
-- on Fin m โ Bool), this is a finite union.MEASURABILITY CONCERNS:
{xs | โ h โ C, ...} is NOT obviously measurable for infinite C. Strategy: decompose via restriction patterns. On any fixed xs, the set of labelings {(h(xs 0), ..., h(xs(m-1))) | h โ C} has at most GF(C,m) โค 2^m elements. So the โh event is a finite union of measurable sets.EmpiricalError is a finite sum of measurable functions, hence measurable.References: SSBD Lemma 4.5, Kakade-Tewari Lecture 19 Lemma 1
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
TrueErrorReal X h c D - EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}On the double sample, the probability that any hypothesis has EmpErr' - EmpErr โฅ ฮต/2 is bounded by GF(C,2m) ยท exp(-mฮตยฒ/8).
Proof strategy (Approach A โ standard exchangeability, 5 steps):
1. EXCHANGEABILITY: Under D^m โ D^m, the 2m draws zโ,...,z_{2m} are iid from D. The joint distribution is invariant under permutations of {1,...,2m}.
Key lemma: P_{D^mโD^m}[event(S,S')] = E_z[P_{split}[event | z]] where z = merged sample and the split is uniformly random among all C(2m,m) ways to partition z into two groups of m.
-- Measure.pi permutation invariance
have pi_perm_invariant : โ (ฯ : Equiv.Perm (Fin (2*m))),
(Measure.pi (fun _ : Fin (2*m) => D)).map (fun z i => z (ฯ i))
= Measure.pi (fun _ : Fin (2*m) => D) := by ...
-- Consequence: the event probability equals the split-averaged probability
have exchangeability :
DoubleSampleMeasure D m {p | โ h โ C, gap(p) โฅ ฮต/2}
= โซ z, SplitMeasure m {vs | โ h โ C, gap(split z vs) โฅ ฮต/2}
โ(Measure.pi (fun _ : Fin (2*m) => D)) := by ...2. CONDITIONING: For fixed merged sample z of 2m points:
-- Number of distinct patterns
have num_patterns : โ (z : MergedSample X m),
Set.ncard {p : Fin (2*m) โ Bool | โ h โ C, โ i, p i = (h (z i) โ c (z i))}
โค GrowthFunction X C (2*m) := by ...3. PER-PATTERN HOEFFDING ON SPLITS: For fixed z and fixed pattern p: Under uniformly random split (S,S') of z into two groups of m: diff(p, split) = (1/m) โ_{iโS'} a_i - (1/m) โ_{iโS} a_i
This is a function of the random partition. By Hoeffding's inequality for sampling without replacement (Serfling 1974): P_split[diff โฅ ฮต/2] โค exp(-mฮตยฒ/8)
Alternative derivation: Hoeffding without replacement from Hoeffding with replacement (iid signs) via coupling. The without-replacement bound is actually TIGHTER (variance reduction), but the with-replacement bound suffices.
-- Per-pattern concentration
have per_pattern_bound : โ (z : MergedSample X m) (a : Fin (2*m) โ โ)
(ha : โ i, a i โ Set.Icc 0 1),
SplitMeasure m {vs | (1/m) * โ i โ second_group vs, a i
- (1/m) * โ i โ first_group vs, a i โฅ ฮต/2}
โค ENNReal.ofReal (Real.exp (-(m : โ) * (ฮต/2)^2 / 2)) := by ...
-- Note: m*(ฮต/2)^2/2 = mฮตยฒ/84. UNION BOUND: P_split[โ pattern: diff โฅ ฮต/2 | z] โค (number of patterns) ยท max_pattern P_split[diff โฅ ฮต/2] โค GF(C,2m) ยท exp(-mฮตยฒ/8)
have union_bound : โ (z : MergedSample X m),
SplitMeasure m {vs | โ h โ C, gap(split z vs, h) โฅ ฮต/2}
โค ENNReal.ofReal (GrowthFunction X C (2*m) * Real.exp (-(m : โ) * ฮต^2 / 8))
:= by ...5. INTEGRATE: P_{D^mโD^m}[event] = E_z[P_split[event|z]] (by step 1) โค E_z[GF(C,2m) ยท exp(-mฮตยฒ/8)] (by step 4, pointwise) = GF(C,2m) ยท exp(-mฮตยฒ/8) (bound is independent of z)
-- The bound is a constant, so integrating gives the same constant
-- (using IsProbabilityMeasure for the 2m-fold product)Infrastructure needed:
Fin.sumFinEquiv : Fin m โ Fin n โ Fin (m + n) (available in Mathlib)mergeSamples / splitMergedSample (defined above)SplitMeasure and ValidSplit (defined above)Measure.pi permutation invariance (to be proved or imported)GrowthFunction on 2m points + sauer_shelah_exp_bound from Rademacher.leanMEASURABILITY CONCERNS:
GrowthFunction X C (2*m) is a natural number (deterministic), no measurability issue.References: SSBD Theorem 6.7, Hoeffding (1963), Serfling (1974)
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
ฮต โค 2 โ
Set.Nonempty C โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
have ฮผ := MeasureTheory.Measure.pi fun x => D;
(ฮผ.prod ฮผ)
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ (Y : Type v) {inst : DecidableEq Y} [inst_1 : DecidableEq Y] (a a_1 : Y),
a = a_1 โ โ (a_2 a_3 : Y), a_2 = a_3 โ zeroOneLoss Y a a_2 = zeroOneLoss Y a_1 a_3The number of distinct restriction patterns of C on any n points is at most GF(C,n). For z : Fin n โ X, define patterns(z) = {p : Fin n โ Bool | โ h โ C, โ i, p i = (h(z i) โ c(z i))}. Then patterns(z).ncard โค GrowthFunction X C n by definition of GrowthFunction.
โ {X : Type u} [MeasurableSpace X] [Infinite X] (C : ConceptClass X Bool) (c : Concept X Bool) (n : โ) (z : Fin n โ X),
{p | โ h โ C, โ (i : Fin n), p i = decide (h (z i) โ c (z i))}.ncard โค GrowthFunction X C nRademacher MGF bound.
โ {m : โ},
0 < m โ
โ (a : Fin m โ โ) (c : โ),
0 โค c โ
(โ (i : Fin m), |a i| โค c) โ
โ (t : โ),
0 โค t โ
1 / โ(Fintype.card (SignVector m)) * โ ฯ, Real.exp (t * (1 / โm * โ i, a i * boolToSign (ฯ i))) โค
Real.exp (t ^ 2 * c ^ 2 / (2 * โm))cosh(x) โค exp(xยฒ/2). Standard sub-Gaussian bound.
โ (x : โ), Real.cosh x โค Real.exp (x ^ 2 / 2)
Generic finite exchangeability bound. Given a measure-preserving family of transformations on a probability space, a NullMeasurableSet S, and a pointwise bound on the sum of preimage indicators, conclude ฮฝ(S) โค B.
โ {ฮฉ : Type u_1} {G : Type u_2} [inst : MeasurableSpace ฮฉ] [inst_1 : Fintype G] [Nonempty G]
{ฮฝ : MeasureTheory.Measure ฮฉ} [MeasureTheory.IsProbabilityMeasure ฮฝ] (T : G โ ฮฉ โ ฮฉ) (S : Set ฮฉ),
(โ (g : G), MeasureTheory.MeasurePreserving (T g) ฮฝ ฮฝ) โ
MeasureTheory.NullMeasurableSet S ฮฝ โ
โ (B : ENNReal), (โ (z : ฮฉ), โ g, (T g โปยน' S).indicator 1 z โค B * โ(Fintype.card G)) โ ฮฝ S โค Bโ {X : Type u} [MeasurableSpace X] (C : ConceptClass X Bool) (v : โ),
0 < v โ
โ (m : โ),
0 < m โ
โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
ฮด < 1 โ
(โ (n : โ), v โค n โ GrowthFunction X C n โค โ i โ Finset.range (v + 1), n.choose i) โ
(16 * Real.exp 1 * (โv + 1) / ฮต ^ 2) ^ (v + 1) / ฮด โค โm โ
4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)) โค ฮด โง 2 * Real.log 2 โค โm * ฮต ^ 2Pure combinatorial inequality: โ_{i=0}^d C(m,i) โค (em/d)^d for d โค m, d โฅ 1.
โ (d m : โ), 0 < d โ d โค m โ โ i โ Finset.range (d + 1), โ(m.choose i) โค (Real.exp 1 * โm / โd) ^ d
Key arithmetic lemma for PAC bound: for t > 0, t^d * exp(-t) โค (d+1)!/t. Follows from exp(t) โฅ t^(d+1)/(d+1)! (partial sum of Taylor series).
โ {d : โ} {t : โ}, 0 < t โ t ^ d * Real.exp (-t) โค โ(d + 1).factorial / tTrivial bound: GrowthFunction โค 2^n for all concept classes. Each restriction to an n-element set yields a function in S โ Bool, and there are at most 2^n such functions.
โ {X : Type u} (C : ConceptClass X Bool) (n : โ), GrowthFunction X C n โค 2 ^ nUpper-tail Hoeffding: for iid Bernoulli(p) draws, the empirical average overshoots the mean by โฅ t with probability โค exp(-2mtยฒ).
This is the mirror of hoeffding_one_sided (which bounds the lower tail). The proof uses the same sub-Gaussian machinery with Z_i = indicator(x_i) - p (instead of p - indicator(x_i)).
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ TrueErrorReal X h c D + t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))One-sided Hoeffding: for iid Bernoulli(p) draws, the empirical average undershoots the mean by โฅ t with probability โค exp(-2mtยฒ).
Proof strategy (3 steps):
1. MGF bound (Hoeffding's lemma): For X โ [0,1] with E[X] = p, E[exp(s(X-p))] โค exp(sยฒ/8).
cosh_le_exp_sq_half infrastructure in Rademacher.lean. have mgf_bound : โ (s : โ), โซ x, Real.exp (s * (indicator x - p)) โD โค Real.exp (s^2 / 8) := by ... 2. Product independence: E[exp(sยทโ(X_i-p))] = โ E[exp(s(X_i-p))] โค exp(msยฒ/8).
MeasureTheory.Measure.pi independence structure.Measure.pi integral factorization for product of functions.fun xs => Real.exp (s * โ i, f (xs i)) is measurable (composition of measurable functions). have product_bound : โ (s : โ), โซ xs, Real.exp (s * โ i, (indicator (xs i) - p)) โMeasure.pi (fun _ => D) โค Real.exp (m * s^2 / 8) := by ... 3. Exponential Markov + optimize: P[โ(X_i-p) โค -mt] = P[exp(-sยทโ(X_i-p)) โฅ exp(smt)] โค exp(-smt + msยฒ/8). Optimize over s: set s = 4t to get โค exp(-2mtยฒ).
have markov_step : โ (s : โ) (hs : 0 < s), Measure.pi (fun _ => D) {xs | โ i, (indicator (xs i) - p) โค -(m : โ) * t} โค ENNReal.ofReal (Real.exp (-(s * m * t) + m * s^2 / 8)) := by ... have optimize : Real.exp (-(4*t * m * t) + m * (4*t)^2 / 8) = Real.exp (-2 * m * t^2) := by ring_nf CAST ISSUES to watch:
m : โ needs cast to โ in the exponent: (m : โ)EmpiricalError returns โ, TrueErrorReal returns โ, good โ no ENNReal gapENNReal, the bound exp(-2mtยฒ) is โโฅ0โ via ENNReal.ofRealReferences: SSBD Lemma B.3, Hoeffding (1963)
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โค TrueErrorReal X h c D - t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))Uniform convergence implies PAC learnability via ERM. The ERM learner (which exists by ermLearn) achieves PAC learning when uniform convergence holds. This is the second half of vcdim_finite_imp_pac.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool), Set.Nonempty C โ HasUniformConvergence X C โ PACLearnable X C
Output is in the hypothesis space
โ {X : Type u} {Y : Type v} (self : BatchLearner X Y) {m : โ} (S : Fin m โ X ร Y), self.learn S โ self.hypothesesAdversarial Rademacher lower bound on shattered sets. For |T| >= 4m^2 + 1, exists D with Rad_m(C,D) >= 1/2.
Proof: D = uniform on T. Product measure = uniform on T^m. On injective samples from T (shattered): EmpRad = 1 (by empRad_eq_one_of_injective_in_shattered). EmpRad โฅ 0 everywhere (by empRad_nonneg). Birthday bound: P[injective m draws from n โฅ 4mยฒ+1 points] โฅ 1 - m(m-1)/(2n) โฅ 7/8 โฅ 1/2. So โซ EmpRad โฅ P[injective] ยท 1 + P[ยฌinjective] ยท 0 โฅ 1/2.
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool) (T : Finset X),
Shatters X C T โ
โ (m : โ),
0 < m โ 4 * m ^ 2 + 1 โค T.card โ โ D, MeasureTheory.IsProbabilityMeasure D โง 1 / 2 โค RademacherComplexity X C D mโ (X : Type u) (C : ConceptClass X Bool) {m : โ}, 0 < m โ โ (xs : Fin m โ X), EmpiricalRademacherComplexity X C xs โค 1โ {X : Type u} (C : ConceptClass X Bool) {m : โ}, m โ 0 โ โ (xs : Fin m โ X), 0 โค EmpiricalRademacherComplexity X C xsRademacher cancellation: ฮฃ_ฯ boolToSign(ฯ i) * f(ฯ) = 0 when f doesn't depend on coordinate i. Proof: the bit-flip involution at coordinate i pairs each ฯ with flipAt i ฯ, negating boolToSign(ฯ i) while preserving f.
โ {m : โ} (i : Fin m) (f : SignVector m โ โ),
(โ (ฯ ฯ' : SignVector m), (โ (k : Fin m), k โ i โ ฯ k = ฯ' k) โ f ฯ = f ฯ') โ โ ฯ, boolToSign (ฯ i) * f ฯ = 0โ {m : โ} (i : Fin m) (ฯ : SignVector m) (k : Fin m), k โ i โ flipAtโ i ฯ k = ฯ kโ {m : โ} (i : Fin m), Function.Involutive (flipAtโ i)โ {m : โ} (i : Fin m) (ฯ : SignVector m), boolToSign (flipAtโ i ฯ i) = -boolToSign (ฯ i)Key combinatorial lemma: injective samples from a shattered set have EmpRad = 1.
โ {X : Type u} [DecidableEq X] (C : ConceptClass X Bool) {m : โ},
0 < m โ
โ (T : Finset X),
Shatters X C T โ
โ (xs : Fin m โ X), Function.Injective xs โ (โ (i : Fin m), xs i โ T) โ EmpiricalRademacherComplexity X C xs = 1Subset of a shattered set is shattered.
โ {X : Type u} (C : ConceptClass X Bool) (T S : Finset X), S โ T โ Shatters X C T โ Shatters X C SOn samples where every labeling is realizable, EmpRad = 1.
For each sign vector ฯ, the hypothesis provides h โ C with h(xs i) = ฯ i, giving corr(h,ฯ,xs) = 1. Since |corr| โค 1, the sSup is exactly 1. Averaging over all ฯ gives EmpRad = (1/2^m)ยท2^mยท1 = 1.
This is the combinatorial core of the NFL Rademacher lower bound: when xs are distinct points from a shattered set, every labeling is realized, so this lemma applies.
โ {X : Type u} (C : ConceptClass X Bool) {m : โ},
0 < m โ
โ (xs : Fin m โ X),
(โ (ฯ : SignVector m), โ h โ C, โ (i : Fin m), h (xs i) = ฯ i) โ EmpiricalRademacherComplexity X C xs = 1โ {X : Type u} {m : โ},
0 < m โ โ (h : Concept X Bool) (ฯ : SignVector m) (xs : Fin m โ X), |rademacherCorrelation h ฯ xs| โค 1โ (bโ bโ : Bool), |boolToSign bโ * boolToSign bโ| โค 1
โ (b : Bool), |boolToSign b| โค 1
โ (b : Bool), |boolToSign b| = 1
When h agrees with ฯ on all sample points, correlation is exactly 1.
โ {X : Type u} {m : โ},
0 < m โ
โ (h : Concept X Bool) (ฯ : SignVector m) (xs : Fin m โ X),
(โ (i : Fin m), h (xs i) = ฯ i) โ rademacherCorrelation h ฯ xs = 1Direction โ: PAC learnability implies finite VCDim.
PROOF ROUTE (via double-sample infrastructure in Generalization.lean): Step 1: Contrapositive โ assume VCDim = โ Step 2: For m = mf(ฮต,ฮด), extract S with |S| = 2m shattered by C (uses WithTop.eq_top_iff_forall_ge, same as vcdim_univ_infinite) Step 3: Construct D = uniform on S (Finset.uniformMeasure?) KUโโ: Mathlib's uniform measure on a finite set โ does MeasureTheory.Measure.count / Finset.card give IsProbabilityMeasure? Step 4: Double-sample trick via GhostSample + symmetrization Step 5: Counting argument on restricted labelings
HC at this joint: Step 3 requires constructing a specific probability measure from a combinatorial object (the shattered set). This is a PโโPโ crossing. UKโ: The construction of the hard distribution is the only non-constructive step. Can it be made constructive? (Related to derandomization in learning.)
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool), PACLearnable X C โ VCDim X C < โค
If VCDim = โค, then C is not PAC learnable. Proof: for any learner L with sample function mf, pick ฮต = 1/4, ฮด = 1/4. Let m = mf(1/4, 1/4). Since VCDim = โค, โ shattered set S with |S| โฅ 2m. Put D = uniform on S. For random labeling, any m-sample learner has expected error โฅ 1/4 on unseen points. This is the core of pac_imp_vcdim_finite (contrapositive direction).
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool), VCDim X C = โค โ ยฌPACLearnable X C
The uniform measure is a probability measure when X is nonempty and finite.
โ (X : Type u) [inst : MeasurableSpace X] [inst_1 : Fintype X] [MeasurableSingletonClass X] (hne : Nonempty X), 0 < Fintype.card X โ MeasureTheory.IsProbabilityMeasure (uniformMeasure X hne)
NFL counting core: for a shattered set T with |T| > 2m, there exists a labeling fโ : โฅT โ Bool and its shattering witness cโ โ C such that the number of samples xs : Fin m โ โฅT where the learner achieves low error (โค |T|/4) is at most half the total number of samples. Proof: double-counting + pigeonhole using per_sample_labeling_bound.
โ {X : Type u} {C : ConceptClass X Bool} {T : Finset X},
Shatters X C T โ
โ {m : โ},
2 * m < T.card โ
โ (L : BatchLearner X Bool),
โ fโ,
โ cโ โ C,
(โ (t : โฅT), cโ โt = fโ t) โง
2 * {xs | {t | cโ โt โ L.learn (fun i => (โ(xs i), cโ โ(xs i))) โt}.card * 4 โค T.card}.card โค
Fintype.card (Fin m โ โฅT)Per-sample labeling bound: for any fixed xs : Fin m โ ฮฑ on a Fintype ฮฑ with 2m < |ฮฑ|, and any function output : (ฮฑ โ Bool) โ (ฮฑ โ Bool) that only depends on the restriction of f to {xs i}, at most half the labelings f : ฮฑ โ Bool have error(f, output(f)) * 4 โค |ฮฑ|.
Proof: pair each f with flip_unseen(f). The pair has complementary disagreements on unseen points, and |unseen| > |ฮฑ|/2, so at most one can have low error.
โ {ฮฑ : Type u_1} [inst : Fintype ฮฑ] [inst_1 : DecidableEq ฮฑ] (m : โ),
2 * m < Fintype.card ฮฑ โ
โ (xs : Fin m โ ฮฑ) (output : (ฮฑ โ Bool) โ ฮฑ โ Bool),
(โ (f f' : ฮฑ โ Bool), (โ (i : Fin m), f (xs i) = f' (xs i)) โ output f = output f') โ
2 * {f | {t | f t โ output f t}.card * 4 โค Fintype.card ฮฑ}.card โค Fintype.card (ฮฑ โ Bool)โ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C],
โ c โ C, Measurable cEvery concept in C is measurable
โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableConceptClass X C],
โ h โ C, Measurable hโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C]
(c : Concept X Bool), Measurable cAll concepts X โ Bool are measurable (for disagreement sets)
โ {X : Type u} {inst : MeasurableSpace X} (C : ConceptClass X Bool) [self : MeasurableConceptClass X C]
(c : Concept X Bool), Measurable cโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C], WellBehavedVC X CUniform convergence bad event is NullMeasurableSet
โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableConceptClass X C],
WellBehavedVC X CA batch learner (PAC paradigm): takes a finite sample, returns a hypothesis.
Type u โ Type v โ Type (max u v)
The learner's hypothesis space
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ HypothesisSpace X YThe learning algorithm: given a sample, produce a hypothesis
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ {m : โ} โ (Fin m โ X ร Y) โ Concept X YA labeled compression scheme with finite side information. This is the object proved to exist by Moran-Yehudayoff (2016, arXiv:1503.06960).
The current CompressionScheme is strictly stronger: it requires reconstruction from the compressed Finset alone (no side information). See Open_NoInfoCompressionStrengthening for that conjecture.
(X : Type u) โ (Y : Type v) โ ConceptClass X Y โ Type (max (max u (u_1 + 1)) v)
The side information type
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ CompressionSchemeWithInfo X Y C โ Type u_1Compression: extract โค kernelSize labeled examples + side information
{X : Type u} โ
{Y : Type v} โ
{C : ConceptClass X Y} โ
(self : CompressionSchemeWithInfo X Y C) โ {m : โ} โ (Fin m โ X ร Y) โ Finset (X ร Y) ร self.InfoSide information is finite
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ (self : CompressionSchemeWithInfo X Y C) โ Fintype self.InfoKernel size bound
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ CompressionSchemeWithInfo X Y C โ โReconstruction: produce hypothesis from compressed subset AND side information
{X : Type u} โ
{Y : Type v} โ {C : ConceptClass X Y} โ (self : CompressionSchemeWithInfo X Y C) โ Finset (X ร Y) โ self.Info โ X โ YTotal size of a compression scheme with side information: kernel size + number of side information states. (The paper uses k + logโ(|I|+1); we use the simpler k + |I| which is an upper bound and avoids importing Real.log.)
{X : Type u} โ {Y : Type v} โ {C : ConceptClass X Y} โ CompressionSchemeWithInfo X Y C โ โFix the hidden Info universe parameter of CompressionSchemeWithInfo to 0. This resolves the universe elaboration obstruction: Fin T โ Finset (Fin K) is Type 0, while CompressionSchemeWithInfo X Bool C with X : Type u infers Info : Type u. Pinning to .{u, 0, 0} allows Type 0 Info directly.
(X : Type u) โ (Y : Type) โ ConceptClass X Y โ Type (max (max u 1) 0)
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
Empirical error: average loss on a finite sample.
(X : Type u) โ (Y : Type v) โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ LossFunction Y โ โ(X : Type u) โ ConceptClass X Bool โ {m : โ} โ (Fin m โ X) โ โA concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
A probability mass function over a finite type. Named FinitePMF to avoid conflict with Mathlib's PMF.
(H : Type u_1) โ [Fintype H] โ Type u_1
{H : Type u_1} โ [inst : Fintype H] โ FinitePMF H โ H โ โ{H : Type u_1} โ [inst : Fintype H] โ FinitePMF H โ PMF HVC dimension of a finite Bool-valued family, computed via the set-system image boolFamilyToFinsetFamily and Mathlib's Finset.vcDim. Declared noncomputable because the underlying vcDim is.
{H : Type u_1} โ [Fintype H] โ [DecidableEq H] โ Finset (H โ Bool) โ โGrowth function (shattering coefficient): ฯ_C(m) = max_{|S|=m} |{c|_S : c โ C}|. For each m-element set S, counts the number of distinct restrictions of C to S, then takes the supremum over all such S.
(X : Type u) โ ConceptClass X Bool โ โ โ โ
Uniform convergence of empirical error to true error over a hypothesis class. This is the property that makes finite VCDim โ PAC learnability work. BPโ connects here: this is ONE of the five characterizations.
M-DefinitionRepair (ฮโโ โ ฮโโ): The mโ must be INDEPENDENT of D and c. That's what "uniform" means โ convergence is uniform over all distributions and all target concepts. The original definition had mโ depending on D and c, making uc_imp_pac unprovable (PACLearnable's mf must be independent of D, c). Repaired: โ mโ is now BEFORE โ D, โ c. This STRENGTHENS the definition (A5-valid: adds content, doesn't simplify).
(X : Type u) โ [MeasurableSpace X] โ HypothesisSpace X Bool โ Prop
A hypothesis space is a set of candidate concepts that the learner searches over. When H = C (the realizable case), every concept in the target class is available. When H โ C or H โ C, we are in the improper/agnostic regime.
Structurally identical to ConceptClass but semantically distinct: ConceptClass is the ground truth collection; HypothesisSpace is what the learner has access to.
Type u โ Type v โ Type (max v u)
Concrete side information for the MY construction: each of the T recovered blocks is represented by the set of kernel positions it uses.
โ โ โ โ Type
A hypothesis h is consistent with labeled sample S.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ PropMWU config: weight vector with positivity proof.
(C : Type u_1) โ [Fintype C] โ Type u_1
Potential = sum of weights.
{C : Type u_1} โ [inst : Fintype C] โ MWUConfig C โ โNormalize config to PMF.
{C : Type u_1} โ [inst : Fintype C] โ [Nonempty C] โ MWUConfig C โ FinitePMF C{C : Type u_1} โ [inst : Fintype C] โ MWUConfig C โ C โ โA concept class with the measure-theoretic regularity needed for PAC theory.
Bundles three conditions: 1. Every concept in C is measurable 2. All concepts are measurable (needed for disagreement set measurability) 3. The UC bad event satisfies NullMeasurableSet (WellBehavedVC)
Condition 3 is the deep one: for uncountable C, the existential {โ h โ C, |TrueErr - EmpErr| โฅ ฮต} is NOT MeasurableSet in general. WellBehavedVC asserts it is NullMeasurableSet, which suffices for integration (lintegral_indicator_oneโ).
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
PAC (Probably Approximately Correct) learning. The central definition of computational learning theory.
Sample space: Fin m โ X with i.i.d. product measure D^m. Labels: derived deterministically from target concept c (realizable case). Error: D-probability of disagreement between learner output and c.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
A proper finite-support learner for a concept class C. This structure captures the existence of a bounded-support ERM with error at most 1/3 for any C-realizable finite distribution. CORRECTED: good_on_support returns Finset X (not Fin k โ X).
(X : Type u) โ ConceptClass X Bool โ Type u
{X : Type u} โ {C : ConceptClass X Bool} โ ProperFiniteSupportLearner X C โ {m : โ} โ (Fin m โ X ร Bool) โ X โ Bool{X : Type u} โ {C : ConceptClass X Bool} โ ProperFiniteSupportLearner X C โ โ(X : Type u) โ [inst : MeasurableSpace X] โ ConceptClass X Bool โ MeasureTheory.Measure X โ โ โ โ
Sample complexity of PAC learning: the minimum number of samples needed to achieve (ฮต,ฮด)-PAC learning. m_C(ฮต,ฮด) = sInf{m | โ L, โ D prob, โ c โ C, D^m{S : error(L(S)) โค ฮต} โฅ 1-ฮด}.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ โ โ โ โ โ
A set S โ X is shattered by concept class C if every labeling of S is realized by some concept in C.
(X : Type u) โ ConceptClass X Bool โ Finset X โ Prop
โ โ Type
True error (0-1 loss, realizable case): D-probability of disagreement. This is what PACLearnable's success event measures.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ ENNReal
True error in โ: for use in bounds involving subtraction/absolute value. COUNTER-1 of TrueError. The toReal bridge loses information when the measure is โค.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ โ
VC dimension of a concept class: the size of the largest shattered set. Returns โโ = WithTop โ.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
A concept class is well-behaved if the ghost gap event is null-measurable. This is the minimal regularity assumption for the symmetrization proof.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Per-point agreement test: for a fixed point x โ Y and concept c, maps hypothesis h to whether h(x) = c(x).
{X : Type u} โ (X โ Bool) โ X โ (HY : Finset (X โ Bool)) โ โฅHY โ BoolThe family of agreement tests over all points in Y.
{X : Type u} โ (X โ Bool) โ Finset X โ (HY : Finset (X โ Bool)) โ Finset (โฅHY โ Bool)Maps a finite family of Bool-valued functions to its image as a family of accepting sets. The set-system view is what Mathlib's Finset.Shatters and Finset.vcDim consume, so this is the entry point from the function-class view to the combinatorial VC machinery.
{H : Type u_1} โ [Fintype H] โ [DecidableEq H] โ Finset (H โ Bool) โ Finset (Finset H)Expected payoff of distribution p against column c in a Boolean game.
{R : Type u_1} โ {C : Type u_2} โ [inst : Fintype R] โ (R โ C โ Bool) โ FinitePMF R โ C โ โExpected value of a Bool-valued test under a finite distribution, via the indicator embedding if f h then 1 else 0. The central quantity of the finite-VC approximation layer: a TV bound on distributions translates to a uniform bound on test expectations via expectation_approx_of_tv.
{H : Type u_1} โ [inst : Fintype H] โ FinitePMF H โ (H โ Bool) โ โConvert Bool labels to ยฑ1 reals. true โฆ 1, false โฆ -1.
Bool โ โ
Bounded subsamples: all subsets of Y with cardinality โค s.
{X : Type u} โ Finset X โ โ โ Finset (Finset X)Decode labels from the kernel. This is exactly the current MY reconstruction convention in your file.
{X : Type u} โ [DecidableEq X] โ Finset (X ร Bool) โ X โ BoolDecode the X-coordinates of a block from kernel positions. This matches the current blockHyp shape.
{X : Type u} โ Finset (X ร Bool) โ {K : โ} โ Finset (Fin K) โ Finset XThe disagreement family: for each h โ C, the test y โฆ decide(h(y) โ c(y)) restricted to Y. Used for the VC approximation step in the proper learner proof.
{X : Type u} โ ConceptClass X Bool โ (X โ Bool) โ (Y : Finset X) โ Finset (โฅY โ Bool)Build FinitePMF from empirical frequencies of a finite sequence.
{ฮฑ : Type u_1} โ [inst : Fintype ฮฑ] โ [DecidableEq ฮฑ] โ {T : โ} โ 0 < T โ (Fin T โ ฮฑ) โ FinitePMF ฮฑEncode a witness set W as the set of kernel positions of the pairs (x, c x). The bound kernel.card โค K is fed into the encoding through the if branch, so the result has the same shape as the current compressCore code.
{X : Type u} โ [DecidableEq X] โ Finset (X ร Bool) โ (X โ Bool) โ (K : โ) โ Finset X โ Finset (Fin K){H : Type u_1} โ (H โ Bool) โ H โ โ โ BoolBit-flip at coordinate i: ฯ โฆ ฯ' where ฯ'(i) = !ฯ(i), ฯ'(k) = ฯ(k) for k โ i.
{m : โ} โ Fin m โ SignVector m โ SignVector mThe hypothesis envelope: the finite set of all possible learner outputs on bounded subsamples of Y, labeled by concept c.
{X : Type u} โ {C : ConceptClass X Bool} โ ProperFiniteSupportLearner X C โ (X โ Bool) โ Finset X โ Finset (X โ Bool)Build a labeled sample from a Finset of points and a concept.
{X : Type u} โ (X โ Bool) โ (Z : Finset X) โ Fin Z.card โ X ร Bool{H : Type u_1} โ Finset (H โ Bool) โ ConceptClass (H โ โ) BoolThe actual final closure helper. Packages the majority-vote construction. If decoded hypotheses agree with reference hypotheses on sample points, and majority of reference hypotheses agree with each label, then majority-vote reconstruction is correct.
{X : Type u} โ
{C : ConceptClass X Bool} โ
(T K : โ) โ
(compressCore : {m : โ} โ (Fin m โ X ร Bool) โ Finset (X ร Bool) ร IncidenceInfo T K) โ
(blockHyp : Finset (X ร Bool) โ IncidenceInfo T K โ Fin T โ X โ Bool) โ
(rowHyp :
{m : โ} โ (S : Fin m โ X ร Bool) โ (โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) โ Fin T โ X โ Bool) โ
0 < T โ
(โ {m : โ} (S : Fin m โ X ร Bool), (compressCore S).1.card โค K) โ
(โ {m : โ} (S : Fin m โ X ร Bool), โ(compressCore S).1 โ Set.range S) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m)
(t : Fin T),
blockHyp (compressCore S).1 (compressCore S).2 t (S i).1 = rowHyp S hreal t (S i).1) โ
(โ {m : โ} (S : Fin m โ X ร Bool) (hreal : โ c โ C, โ (i : Fin m), c (S i).1 = (S i).2) (i : Fin m),
(โ t, if rowHyp S hreal t (S i).1 = (S i).2 then 1 else 0) / โT > 1 / 2) โ
CompressionSchemeWithInfo0 X Bool CThe MWU config after T steps.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ โ โ MWUConfig CCount how many rounds hit a fixed column, aligned to the recursion of mwuRun.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ (โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ โ โ C โ โInitial config: all weights = 1.
(C : Type u_1) โ [inst : Fintype C] โ MWUConfig C
The MWU row sequence after T steps.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ (T : โ) โ Fin T โ RMWU run: iterate T steps, returning final config and row sequence.
{R : Type u_1} โ
{C : Type u_2} โ
[Fintype R] โ
[inst : Fintype C] โ
[Nonempty C] โ
(M : R โ C โ Bool) โ
(ฮท : โ) โ
ฮท < 1 โ
(v : โ) โ
(โ (q : FinitePMF C), โ r, v โค โ c, q.prob c * if M r c = true then 1 else 0) โ
(T : โ) โ MWUConfig C ร (Fin T โ R)One MWU update step on weights.
{C : Type u_1} โ [inst : Fintype C] โ {R : Type u_2} โ (R โ C โ Bool) โ (ฮท : โ) โ ฮท < 1 โ MWUConfig C โ R โ MWUConfig CExtract the domain points from a labeled sample.
{X : Type u} โ {m : โ} โ (Fin m โ X ร Bool) โ Finset X{X : Type u} โ {m : โ} โ Concept X Bool โ SignVector m โ (Fin m โ X) โ โWeighted error of hypothesis h vs concept c over a FinitePMF on Y.
{X : Type u} โ (Y : Finset X) โ FinitePMF โฅY โ (X โ Bool) โ (X โ Bool) โ โUniform probability measure on a Fintype: (1/|X|) ยท count. This gives each point probability 1/|X|. Requires |X| > 0 (nonempty).
(X : Type u) โ [inst : MeasurableSpace X] โ [Fintype X] โ Nonempty X โ MeasureTheory.Measure X
Uniform PMF over a nonempty Fintype.
(C : Type u_1) โ [inst : Fintype C] โ [Nonempty C] โ FinitePMF C
The 0-1 loss for classification.
(Y : Type v) โ [DecidableEq Y] โ LossFunction Y
The fundamental theorem of statistical learning. For a measurable concept class, finite VC dimension, eventually polynomial growth, and PAC learnability are mutually equivalent.
โ {X : Type u} [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool)
[MeasurableConceptClass X C],
(VCDim X C < โค โ PACLearnable X C) โง
(VCDim X C < โค โ โ K d, โแถ (m : โ) in Filter.atTop, โ(GrowthFunction X C m) โค K * โm ^ d)โ {X : Type u} [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool)
[MeasurableConceptClass X C],
(VCDim X C < โค โ PACLearnable X C) โง
(VCDim X C < โค โ โ K d, โแถ (m : โ) in Filter.atTop, โ(GrowthFunction X C m) โค K * โm ^ d)Finite VC dimension โบ PAC learnability. The measure-theoretic half: VCDim X C < โค exactly when C is PAC learnable. This is the kernel's vc_characterization, surfaced for the independent module; it requires the domain's measurability structure.
โ {X : Type u} [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool)
[MeasurableConceptClass X C], VCDim X C < โค โ PACLearnable X CVC characterization: C is PAC-learnable iff VCDim(C) < โ.
PROOF DECOMPOSITION: This theorem factors through the two directions above: โ : vcdim_finite_imp_uc + uc_imp_pac (in Generalization.lean) โ : pac_imp_vcdim_finite (contrapositive via double-sample)
HC at this joint: The โ direction crosses from combinatorics (VCDim, GrowthFunction) to measure theory (Measure.pi, TrueError). The โ direction crosses from measure theory back to combinatorics. Both crossings have HC > 0.
UKโ: The โ hides an ASYMMETRY: the โ proof is constructive (produces ERM), while the โ proof is non-constructive (produces hard distribution).
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool) [MeasurableConceptClass X C], PACLearnable X C โ VCDim X C < โค
Direction โ: finite VCDim implies PAC learnability.
PROOF ROUTE (via new infrastructure in Generalization.lean): Step 1: VCDim < โ โ HasUniformConvergence (vcdim_finite_imp_uc) Sub-step 1a: Sauer-Shelah gives GrowthFunction bound Sub-step 1b: Symmetrization reduces UC to growth function counting Sub-step 1c: Concentration inequality closes the bound Step 2: HasUniformConvergence โ PACLearnable (uc_imp_pac) Sub-step 2a: Construct ERM learner Sub-step 2b: ERM is consistent in realizable case Sub-step 2c: Consistent + UC โ low TrueError
KUโโ: C.Nonempty is needed for ERM but not stated as hypothesis. If C = โ , then PACLearnable is vacuously true (โ c โ C, ... is vacuous). But ERM needs a fallback hypothesis from C. Is this a genuine gap or does the empty case work out vacuously?
Counterdefinition (COUNTER-4): If the ERM approach fails for computational reasons (ERM is noncomputable, and we need a computable learner for computational learning theory), swap to the compression-based proof: VCDim < โ โ finite compression scheme (Moran-Yehudayoff 2016) โ compression scheme learner is PAC. Swap condition: When proving COMPUTATIONAL PAC learnability (polynomial time).
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool), VCDim X C < โค โ โ [MeasurableConceptClass X C], PACLearnable X C
Finite VCDim implies uniform convergence. Proof: VCDim < โ โ UC.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool),
VCDim X C < โค โ
(โ h โ C, Measurable h) โ (โ (c : Concept X Bool), Measurable c) โ WellBehavedVC X C โ HasUniformConvergence X CUC bad-event bound: for m โฅ mโ(v,ฮต,ฮด), the probability of the bad event (โ h with |TrueErr-EmpErr| โฅ ฮต) is at most ฮด. Composes symmetrization_uc_bound with growth_exp_le_delta.
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
ฮด < 1 โ
โ (v : โ),
0 < v โ
(โ (n : โ), v โค n โ GrowthFunction X C n โค โ i โ Finset.range (v + 1), n.choose i) โ
(16 * Real.exp 1 * (โv + 1) / ฮต ^ 2) ^ (v + 1) / ฮด โค โm โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
|TrueErrorReal X h c D -
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮต} โค
ENNReal.ofReal ฮดThe symmetrization uniform convergence bound: two-sided version. P[โhโC: |TrueErr-EmpErr| โฅ ฮต] โค 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8).
Proof strategy (4 steps):
1. Decompose absolute value: |TrueErr - EmpErr| โฅ ฮต โ (TrueErr - EmpErr โฅ ฮต) โจ (EmpErr - TrueErr โฅ ฮต)
have abs_decomp : โ (a b : โ),
|a - b| โฅ ฮต โ a - b โฅ ฮต โจ b - a โฅ ฮต := by
intro a b; constructor
ยท intro h; by_cases h' : a - b โฅ ฮต
ยท exact Or.inl h'
ยท exact Or.inr (by linarith [abs_sub_comm a b, le_abs_self (a - b)])
ยท intro h; cases h with
| inl h => exact le_trans (le_of_eq (abs_of_nonneg (by linarith))) (by linarith)
| inr h => exact le_trans (le_of_eq (abs_of_nonpos (by linarith) โธ ...)) ...2. Upper tail: P[โhโC: TrueErr-EmpErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step + double_sample_pattern_bound.3. Lower tail: P[โhโC: EmpErr-TrueErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step to the event EmpErr-TrueErr โฅ ฮต and bound the double-sample event {EmpErr_S - EmpErr_{S'} โฅ ฮต/2}. have swap_symmetry : DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S) - EmpErr(S') โฅ ฮต/2} = DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S') - EmpErr(S) โฅ ฮต/2} := Measure.prod_swap ... 4. Union bound: P[|gap| โฅ ฮต] โค P[gap โฅ ฮต] + P[gap โค -ฮต] โค 2ยทGFยทexp(...) + 2ยทGFยทexp(...) = 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
-- Uses: MeasureTheory.measure_union_le for the union of two events
-- CAST: 2 * X + 2 * X = 4 * X in ENNReal (need ENNReal.add_mul or similar)References: SSBD Theorem 6.7, Kakade-Tewari Lecture 19
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
|TrueErrorReal X h c D -
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮต} โค
ENNReal.ofReal (4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))Symmetrization step for the lower tail: P[โh: EmpErr-TrueErr โฅ ฮต] โค 2ยทP_{double}[โh: EmpErr_S-EmpErr_{S'} โฅ ฮต/2].
Mirror of symmetrization_step for the opposite direction. Uses hoeffding_one_sided_upper instead of hoeffding_one_sided.
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) - TrueErrorReal X h c D โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}Symmetrization: the probability of a large gap TrueErr-EmpErr is at most twice the probability of a large gap EmpErr'-EmpErr on the double sample.
Proof strategy (6 steps):
1. Witness selection: For S in the bad event, โh* โ C with TrueErr(h) - EmpErr_S(h) โฅ ฮต.
-- In the bad event set, extract h* by classical choice
have h_witness : โ xs โ bad_event, โ h* โ C,
TrueErrorReal X h* c D - EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โฅ ฮต2. Ghost sample mean: E_{S'}[EmpErr_{S'}(h)] = TrueErr(h) โฅ EmpErr_S(h*) + ฮต.
MeasureTheory.integral_pi to compute E[EmpErr] over product measure. have expected_emp_err : โ h* : Concept X Bool, โซ xs, EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โ(Measure.pi (fun _ : Fin m => D)) = TrueErrorReal X h* c D := by ... 3. Hoeffding on ghost sample: P_{S'}[EmpErr_{S'}(h) < TrueErr(h) - ฮต/2] โค exp(-mฮตยฒ/2).
hoeffding_one_sided with t = ฮต/2.hm_large hypothesis ensures exp(-mฮตยฒ/2) < 1/2: 2ยทln2 โค mฮตยฒ โน mฮตยฒ/2 โฅ ln2 โน exp(-mฮตยฒ/2) โค 1/2. have hoeffding_ghost : โ h* โ C, Measure.pi (fun _ : Fin m => D) {xs' | EmpiricalError X Bool h* (sample xs') (zeroOneLoss Bool) < TrueErrorReal X h* c D - ฮต/2} โค ENNReal.ofReal (Real.exp (-m * (ฮต/2)^2 * 2)) := by intro h* _; exact hoeffding_one_sided D h* c m hm (ฮต/2) (by linarith) (by ...) (by ...) 4. Complementary probability: P_{S'}[EmpErr_{S'}(h) - EmpErr_S(h) โฅ ฮต/2] โฅ 1/2.
5. Conditional to unconditional: The witness h* from step 1 also witnesses the double-sample event โhโC: EmpErr'-EmpErr โฅ ฮต/2. So: P_{S'}[double event | S bad] โฅ 1/2.
have conditional_bound : โ xs โ bad_event,
Measure.pi (fun _ : Fin m => D)
{xs' | โ h โ C, EmpiricalError ... xs' - EmpiricalError ... xs โฅ ฮต/2}
โฅ ENNReal.ofReal (1/2) := by ...6. Fubini integration: By Measure.prod_apply and Fubini: P_{S,S'}[double event] = โซ_S P_{S'}[double event | S] โฅ (1/2) ยท P_S[bad event] โน P_S[bad event] โค 2 ยท P_{S,S'}[double event].
-- Uses: MeasureTheory.Measure.prod_apply or lintegral_prod
-- MEASURABILITY: the double-sample event is measurable as a finite union
-- of sets of the form {(xs,xs') | EmpErr'(h) - EmpErr(h) โฅ ฮต/2} for h โ C.
-- Since C may be infinite, measurability requires care: the sup over h
-- must be shown to be measurable. For finite restriction patterns (โค 2^m
-- on Fin m โ Bool), this is a finite union.MEASURABILITY CONCERNS:
{xs | โ h โ C, ...} is NOT obviously measurable for infinite C. Strategy: decompose via restriction patterns. On any fixed xs, the set of labelings {(h(xs 0), ..., h(xs(m-1))) | h โ C} has at most GF(C,m) โค 2^m elements. So the โh event is a finite union of measurable sets.EmpiricalError is a finite sum of measurable functions, hence measurable.References: SSBD Lemma 4.5, Kakade-Tewari Lecture 19 Lemma 1
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
TrueErrorReal X h c D - EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}On the double sample, the probability that any hypothesis has EmpErr' - EmpErr โฅ ฮต/2 is bounded by GF(C,2m) ยท exp(-mฮตยฒ/8).
Proof strategy (Approach A โ standard exchangeability, 5 steps):
1. EXCHANGEABILITY: Under D^m โ D^m, the 2m draws zโ,...,z_{2m} are iid from D. The joint distribution is invariant under permutations of {1,...,2m}.
Key lemma: P_{D^mโD^m}[event(S,S')] = E_z[P_{split}[event | z]] where z = merged sample and the split is uniformly random among all C(2m,m) ways to partition z into two groups of m.
-- Measure.pi permutation invariance
have pi_perm_invariant : โ (ฯ : Equiv.Perm (Fin (2*m))),
(Measure.pi (fun _ : Fin (2*m) => D)).map (fun z i => z (ฯ i))
= Measure.pi (fun _ : Fin (2*m) => D) := by ...
-- Consequence: the event probability equals the split-averaged probability
have exchangeability :
DoubleSampleMeasure D m {p | โ h โ C, gap(p) โฅ ฮต/2}
= โซ z, SplitMeasure m {vs | โ h โ C, gap(split z vs) โฅ ฮต/2}
โ(Measure.pi (fun _ : Fin (2*m) => D)) := by ...2. CONDITIONING: For fixed merged sample z of 2m points:
-- Number of distinct patterns
have num_patterns : โ (z : MergedSample X m),
Set.ncard {p : Fin (2*m) โ Bool | โ h โ C, โ i, p i = (h (z i) โ c (z i))}
โค GrowthFunction X C (2*m) := by ...3. PER-PATTERN HOEFFDING ON SPLITS: For fixed z and fixed pattern p: Under uniformly random split (S,S') of z into two groups of m: diff(p, split) = (1/m) โ_{iโS'} a_i - (1/m) โ_{iโS} a_i
This is a function of the random partition. By Hoeffding's inequality for sampling without replacement (Serfling 1974): P_split[diff โฅ ฮต/2] โค exp(-mฮตยฒ/8)
Alternative derivation: Hoeffding without replacement from Hoeffding with replacement (iid signs) via coupling. The without-replacement bound is actually TIGHTER (variance reduction), but the with-replacement bound suffices.
-- Per-pattern concentration
have per_pattern_bound : โ (z : MergedSample X m) (a : Fin (2*m) โ โ)
(ha : โ i, a i โ Set.Icc 0 1),
SplitMeasure m {vs | (1/m) * โ i โ second_group vs, a i
- (1/m) * โ i โ first_group vs, a i โฅ ฮต/2}
โค ENNReal.ofReal (Real.exp (-(m : โ) * (ฮต/2)^2 / 2)) := by ...
-- Note: m*(ฮต/2)^2/2 = mฮตยฒ/84. UNION BOUND: P_split[โ pattern: diff โฅ ฮต/2 | z] โค (number of patterns) ยท max_pattern P_split[diff โฅ ฮต/2] โค GF(C,2m) ยท exp(-mฮตยฒ/8)
have union_bound : โ (z : MergedSample X m),
SplitMeasure m {vs | โ h โ C, gap(split z vs, h) โฅ ฮต/2}
โค ENNReal.ofReal (GrowthFunction X C (2*m) * Real.exp (-(m : โ) * ฮต^2 / 8))
:= by ...5. INTEGRATE: P_{D^mโD^m}[event] = E_z[P_split[event|z]] (by step 1) โค E_z[GF(C,2m) ยท exp(-mฮตยฒ/8)] (by step 4, pointwise) = GF(C,2m) ยท exp(-mฮตยฒ/8) (bound is independent of z)
-- The bound is a constant, so integrating gives the same constant
-- (using IsProbabilityMeasure for the 2m-fold product)Infrastructure needed:
Fin.sumFinEquiv : Fin m โ Fin n โ Fin (m + n) (available in Mathlib)mergeSamples / splitMergedSample (defined above)SplitMeasure and ValidSplit (defined above)Measure.pi permutation invariance (to be proved or imported)GrowthFunction on 2m points + sauer_shelah_exp_bound from Rademacher.leanMEASURABILITY CONCERNS:
GrowthFunction X C (2*m) is a natural number (deterministic), no measurability issue.References: SSBD Theorem 6.7, Hoeffding (1963), Serfling (1974)
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
ฮต โค 2 โ
Set.Nonempty C โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
have ฮผ := MeasureTheory.Measure.pi fun x => D;
(ฮผ.prod ฮผ)
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ (Y : Type v) {inst : DecidableEq Y} [inst_1 : DecidableEq Y] (a a_1 : Y),
a = a_1 โ โ (a_2 a_3 : Y), a_2 = a_3 โ zeroOneLoss Y a a_2 = zeroOneLoss Y a_1 a_3The number of distinct restriction patterns of C on any n points is at most GF(C,n). For z : Fin n โ X, define patterns(z) = {p : Fin n โ Bool | โ h โ C, โ i, p i = (h(z i) โ c(z i))}. Then patterns(z).ncard โค GrowthFunction X C n by definition of GrowthFunction.
โ {X : Type u} [MeasurableSpace X] [Infinite X] (C : ConceptClass X Bool) (c : Concept X Bool) (n : โ) (z : Fin n โ X),
{p | โ h โ C, โ (i : Fin n), p i = decide (h (z i) โ c (z i))}.ncard โค GrowthFunction X C nRademacher MGF bound.
โ {m : โ},
0 < m โ
โ (a : Fin m โ โ) (c : โ),
0 โค c โ
(โ (i : Fin m), |a i| โค c) โ
โ (t : โ),
0 โค t โ
1 / โ(Fintype.card (SignVector m)) * โ ฯ, Real.exp (t * (1 / โm * โ i, a i * boolToSign (ฯ i))) โค
Real.exp (t ^ 2 * c ^ 2 / (2 * โm))cosh(x) โค exp(xยฒ/2). Standard sub-Gaussian bound.
โ (x : โ), Real.cosh x โค Real.exp (x ^ 2 / 2)
Generic finite exchangeability bound. Given a measure-preserving family of transformations on a probability space, a NullMeasurableSet S, and a pointwise bound on the sum of preimage indicators, conclude ฮฝ(S) โค B.
โ {ฮฉ : Type u_1} {G : Type u_2} [inst : MeasurableSpace ฮฉ] [inst_1 : Fintype G] [Nonempty G]
{ฮฝ : MeasureTheory.Measure ฮฉ} [MeasureTheory.IsProbabilityMeasure ฮฝ] (T : G โ ฮฉ โ ฮฉ) (S : Set ฮฉ),
(โ (g : G), MeasureTheory.MeasurePreserving (T g) ฮฝ ฮฝ) โ
MeasureTheory.NullMeasurableSet S ฮฝ โ
โ (B : ENNReal), (โ (z : ฮฉ), โ g, (T g โปยน' S).indicator 1 z โค B * โ(Fintype.card G)) โ ฮฝ S โค Bโ {X : Type u} [MeasurableSpace X] (C : ConceptClass X Bool) (v : โ),
0 < v โ
โ (m : โ),
0 < m โ
โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
ฮด < 1 โ
(โ (n : โ), v โค n โ GrowthFunction X C n โค โ i โ Finset.range (v + 1), n.choose i) โ
(16 * Real.exp 1 * (โv + 1) / ฮต ^ 2) ^ (v + 1) / ฮด โค โm โ
4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)) โค ฮด โง 2 * Real.log 2 โค โm * ฮต ^ 2Pure combinatorial inequality: โ_{i=0}^d C(m,i) โค (em/d)^d for d โค m, d โฅ 1.
โ (d m : โ), 0 < d โ d โค m โ โ i โ Finset.range (d + 1), โ(m.choose i) โค (Real.exp 1 * โm / โd) ^ d
Key arithmetic lemma for PAC bound: for t > 0, t^d * exp(-t) โค (d+1)!/t. Follows from exp(t) โฅ t^(d+1)/(d+1)! (partial sum of Taylor series).
โ {d : โ} {t : โ}, 0 < t โ t ^ d * Real.exp (-t) โค โ(d + 1).factorial / tTrivial bound: GrowthFunction โค 2^n for all concept classes. Each restriction to an n-element set yields a function in S โ Bool, and there are at most 2^n such functions.
โ {X : Type u} (C : ConceptClass X Bool) (n : โ), GrowthFunction X C n โค 2 ^ nUpper-tail Hoeffding: for iid Bernoulli(p) draws, the empirical average overshoots the mean by โฅ t with probability โค exp(-2mtยฒ).
This is the mirror of hoeffding_one_sided (which bounds the lower tail). The proof uses the same sub-Gaussian machinery with Z_i = indicator(x_i) - p (instead of p - indicator(x_i)).
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ TrueErrorReal X h c D + t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))One-sided Hoeffding: for iid Bernoulli(p) draws, the empirical average undershoots the mean by โฅ t with probability โค exp(-2mtยฒ).
Proof strategy (3 steps):
1. MGF bound (Hoeffding's lemma): For X โ [0,1] with E[X] = p, E[exp(s(X-p))] โค exp(sยฒ/8).
cosh_le_exp_sq_half infrastructure in Rademacher.lean. have mgf_bound : โ (s : โ), โซ x, Real.exp (s * (indicator x - p)) โD โค Real.exp (s^2 / 8) := by ... 2. Product independence: E[exp(sยทโ(X_i-p))] = โ E[exp(s(X_i-p))] โค exp(msยฒ/8).
MeasureTheory.Measure.pi independence structure.Measure.pi integral factorization for product of functions.fun xs => Real.exp (s * โ i, f (xs i)) is measurable (composition of measurable functions). have product_bound : โ (s : โ), โซ xs, Real.exp (s * โ i, (indicator (xs i) - p)) โMeasure.pi (fun _ => D) โค Real.exp (m * s^2 / 8) := by ... 3. Exponential Markov + optimize: P[โ(X_i-p) โค -mt] = P[exp(-sยทโ(X_i-p)) โฅ exp(smt)] โค exp(-smt + msยฒ/8). Optimize over s: set s = 4t to get โค exp(-2mtยฒ).
have markov_step : โ (s : โ) (hs : 0 < s), Measure.pi (fun _ => D) {xs | โ i, (indicator (xs i) - p) โค -(m : โ) * t} โค ENNReal.ofReal (Real.exp (-(s * m * t) + m * s^2 / 8)) := by ... have optimize : Real.exp (-(4*t * m * t) + m * (4*t)^2 / 8) = Real.exp (-2 * m * t^2) := by ring_nf CAST ISSUES to watch:
m : โ needs cast to โ in the exponent: (m : โ)EmpiricalError returns โ, TrueErrorReal returns โ, good โ no ENNReal gapENNReal, the bound exp(-2mtยฒ) is โโฅ0โ via ENNReal.ofRealReferences: SSBD Lemma B.3, Hoeffding (1963)
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โค TrueErrorReal X h c D - t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))Uniform convergence implies PAC learnability via ERM. The ERM learner (which exists by ermLearn) achieves PAC learning when uniform convergence holds. This is the second half of vcdim_finite_imp_pac.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool), Set.Nonempty C โ HasUniformConvergence X C โ PACLearnable X C
Output is in the hypothesis space
โ {X : Type u} {Y : Type v} (self : BatchLearner X Y) {m : โ} (S : Fin m โ X ร Y), self.learn S โ self.hypothesesโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C],
โ c โ C, Measurable cEvery concept in C is measurable
โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableConceptClass X C],
โ h โ C, Measurable hโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C]
(c : Concept X Bool), Measurable cAll concepts X โ Bool are measurable (for disagreement sets)
โ {X : Type u} {inst : MeasurableSpace X} (C : ConceptClass X Bool) [self : MeasurableConceptClass X C]
(c : Concept X Bool), Measurable cโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C], WellBehavedVC X CUniform convergence bad event is NullMeasurableSet
โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableConceptClass X C],
WellBehavedVC X CDirection โ: PAC learnability implies finite VCDim.
PROOF ROUTE (via double-sample infrastructure in Generalization.lean): Step 1: Contrapositive โ assume VCDim = โ Step 2: For m = mf(ฮต,ฮด), extract S with |S| = 2m shattered by C (uses WithTop.eq_top_iff_forall_ge, same as vcdim_univ_infinite) Step 3: Construct D = uniform on S (Finset.uniformMeasure?) KUโโ: Mathlib's uniform measure on a finite set โ does MeasureTheory.Measure.count / Finset.card give IsProbabilityMeasure? Step 4: Double-sample trick via GhostSample + symmetrization Step 5: Counting argument on restricted labelings
HC at this joint: Step 3 requires constructing a specific probability measure from a combinatorial object (the shattered set). This is a PโโPโ crossing. UKโ: The construction of the hard distribution is the only non-constructive step. Can it be made constructive? (Related to derandomization in learning.)
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool), PACLearnable X C โ VCDim X C < โค
If VCDim = โค, then C is not PAC learnable. Proof: for any learner L with sample function mf, pick ฮต = 1/4, ฮด = 1/4. Let m = mf(1/4, 1/4). Since VCDim = โค, โ shattered set S with |S| โฅ 2m. Put D = uniform on S. For random labeling, any m-sample learner has expected error โฅ 1/4 on unseen points. This is the core of pac_imp_vcdim_finite (contrapositive direction).
โ (X : Type u) [inst : MeasurableSpace X] [MeasurableSingletonClass X] (C : ConceptClass X Bool), VCDim X C = โค โ ยฌPACLearnable X C
The uniform measure is a probability measure when X is nonempty and finite.
โ (X : Type u) [inst : MeasurableSpace X] [inst_1 : Fintype X] [MeasurableSingletonClass X] (hne : Nonempty X), 0 < Fintype.card X โ MeasureTheory.IsProbabilityMeasure (uniformMeasure X hne)
NFL counting core: for a shattered set T with |T| > 2m, there exists a labeling fโ : โฅT โ Bool and its shattering witness cโ โ C such that the number of samples xs : Fin m โ โฅT where the learner achieves low error (โค |T|/4) is at most half the total number of samples. Proof: double-counting + pigeonhole using per_sample_labeling_bound.
โ {X : Type u} {C : ConceptClass X Bool} {T : Finset X},
Shatters X C T โ
โ {m : โ},
2 * m < T.card โ
โ (L : BatchLearner X Bool),
โ fโ,
โ cโ โ C,
(โ (t : โฅT), cโ โt = fโ t) โง
2 * {xs | {t | cโ โt โ L.learn (fun i => (โ(xs i), cโ โ(xs i))) โt}.card * 4 โค T.card}.card โค
Fintype.card (Fin m โ โฅT)Per-sample labeling bound: for any fixed xs : Fin m โ ฮฑ on a Fintype ฮฑ with 2m < |ฮฑ|, and any function output : (ฮฑ โ Bool) โ (ฮฑ โ Bool) that only depends on the restriction of f to {xs i}, at most half the labelings f : ฮฑ โ Bool have error(f, output(f)) * 4 โค |ฮฑ|.
Proof: pair each f with flip_unseen(f). The pair has complementary disagreements on unseen points, and |unseen| > |ฮฑ|/2, so at most one can have low error.
โ {ฮฑ : Type u_1} [inst : Fintype ฮฑ] [inst_1 : DecidableEq ฮฑ] (m : โ),
2 * m < Fintype.card ฮฑ โ
โ (xs : Fin m โ ฮฑ) (output : (ฮฑ โ Bool) โ ฮฑ โ Bool),
(โ (f f' : ฮฑ โ Bool), (โ (i : Fin m), f (xs i) = f' (xs i)) โ output f = output f') โ
2 * {f | {t | f t โ output f t}.card * 4 โค Fintype.card ฮฑ}.card โค Fintype.card (ฮฑ โ Bool)Finite VC dimension โบ eventually polynomial growth. The purely combinatorial half of the fundamental theorem: VCDim X C < โค exactly when the growth function is eventually bounded by a polynomial.
โ {X : Type u} {C : ConceptClass X Bool},
VCDim X C < โค โ โ K d, โแถ (m : โ) in Filter.atTop, โ(GrowthFunction X C m) โค K * โm ^ dPolynomial growth implies finite VC dimension. If GrowthFunction X C m โค K ยท m ^ d for all large m, then VCDim X C < โค: an infinite VC dimension produces shattered sets of every size, forcing 2 ^ m โค K ยท m ^ d for arbitrarily large m, which is impossible. The hypothesis is the eventually filter (the โ m form is unsatisfiable at m = 0); Sauer-Shelah supplies it for m โฅ d.
โ {X : Type u} {C : ConceptClass X Bool} (K : โ) (d : โ),
(โแถ (m : โ) in Filter.atTop, โ(GrowthFunction X C m) โค K * โm ^ d) โ VCDim X C < โคAn infinite VC dimension forces the growth function to be at least 2 ^ m at every size: it shatters sets of every size, by iSupโ_eq_top and downward closure.
โ {X : Type u} {C : ConceptClass X Bool}, VCDim X C = โค โ โ (m : โ), 2 ^ m โค GrowthFunction X C mShattering is closed under restriction of the sample: if C shatters S and T โ S, then C shatters T. Any labelling of T extends to a labelling of S, which is realized.
โ {X : Type u} {C : ConceptClass X Bool} {S T : Finset X}, Shatters X C S โ T โ S โ Shatters X C TA shattered sample of size m forces the growth function at m to be at least 2 ^ m.
โ {X : Type u} {C : ConceptClass X Bool} {S : Finset X}, Shatters X C S โ 2 ^ S.card โค GrowthFunction X C S.cardEach restriction count is at most the growth function at the matching sample size.
โ {X : Type u} (C : ConceptClass X Bool) {S : Finset X} {m : โ},
S.card = m โ (restrictionSet C S).ncard โค GrowthFunction X C mโ {X : Type u} (C : ConceptClass X Bool) (S : Finset X), (restrictionSet C S).ncard โค 2 ^ S.cardโ {X : Type u} (C : ConceptClass X Bool) (m : โ),
GrowthFunction X C m = sSup (Set.range fun S => (restrictionSet C โS).ncard)A shattered sample realizes every labelling, so its restriction-pattern set is everything.
โ {X : Type u} {C : ConceptClass X Bool} {S : Finset X}, Shatters X C S โ restrictionSet C S = Set.univThe exponential-beats-polynomial crux. 2 ^ m โค K ยท m ^ d cannot hold for arbitrarily large m, since m ^ d = o(2 ^ m). Stated with the eventually filter so it is applicable (the โ m form is unsatisfiable at m = 0 for d โฅ 1).
โ (K : โ) (d : โ), (โแถ (m : โ) in Filter.atTop, 2 ^ m โค K * โm ^ d) โ False
Finite VC dimension implies eventual polynomial growth. If VCDim X C < โค then the growth function is eventually bounded by a polynomial K ยท m ^ d.
โ {X : Type u} {C : ConceptClass X Bool},
VCDim X C < โค โ โ K d, โแถ (m : โ) in Filter.atTop, โ(GrowthFunction X C m) โค K * โm ^ dVCDim < โค โ growth function polynomially bounded by partial binomial sum. Forward direction of fundamental_theorem conjunct 5. Uses Sauer-Shelah: GrowthFunction(m) โค โ_{iโคd} C(m,i) where d = VCDim.
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ d, โ (m : โ), d โค m โ GrowthFunction X C m โค โ i โ Finset.range (d + 1), m.choose i
A batch learner (PAC paradigm): takes a finite sample, returns a hypothesis.
Type u โ Type v โ Type (max u v)
The learner's hypothesis space
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ HypothesisSpace X YThe learning algorithm: given a sample, produce a hypothesis
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ {m : โ} โ (Fin m โ X ร Y) โ Concept X YA concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
Empirical error: average loss on a finite sample.
(X : Type u) โ (Y : Type v) โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ LossFunction Y โ โA concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
Growth function (shattering coefficient): ฯ_C(m) = max_{|S|=m} |{c|_S : c โ C}|. For each m-element set S, counts the number of distinct restrictions of C to S, then takes the supremum over all such S.
(X : Type u) โ ConceptClass X Bool โ โ โ โ
Uniform convergence of empirical error to true error over a hypothesis class. This is the property that makes finite VCDim โ PAC learnability work. BPโ connects here: this is ONE of the five characterizations.
M-DefinitionRepair (ฮโโ โ ฮโโ): The mโ must be INDEPENDENT of D and c. That's what "uniform" means โ convergence is uniform over all distributions and all target concepts. The original definition had mโ depending on D and c, making uc_imp_pac unprovable (PACLearnable's mf must be independent of D, c). Repaired: โ mโ is now BEFORE โ D, โ c. This STRENGTHENS the definition (A5-valid: adds content, doesn't simplify).
(X : Type u) โ [MeasurableSpace X] โ HypothesisSpace X Bool โ Prop
A hypothesis space is a set of candidate concepts that the learner searches over. When H = C (the realizable case), every concept in the target class is available. When H โ C or H โ C, we are in the improper/agnostic regime.
Structurally identical to ConceptClass but semantically distinct: ConceptClass is the ground truth collection; HypothesisSpace is what the learner has access to.
Type u โ Type v โ Type (max v u)
A hypothesis h is consistent with labeled sample S.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ PropA concept class with the measure-theoretic regularity needed for PAC theory.
Bundles three conditions: 1. Every concept in C is measurable 2. All concepts are measurable (needed for disagreement set measurability) 3. The UC bad event satisfies NullMeasurableSet (WellBehavedVC)
Condition 3 is the deep one: for uncountable C, the existential {โ h โ C, |TrueErr - EmpErr| โฅ ฮต} is NOT MeasurableSet in general. WellBehavedVC asserts it is NullMeasurableSet, which suffices for integration (lintegral_indicator_oneโ).
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
PAC (Probably Approximately Correct) learning. The central definition of computational learning theory.
Sample space: Fin m โ X with i.i.d. product measure D^m. Labels: derived deterministically from target concept c (realizable case). Error: D-probability of disagreement between learner output and c.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
A set S โ X is shattered by concept class C if every labeling of S is realized by some concept in C.
(X : Type u) โ ConceptClass X Bool โ Finset X โ Prop
โ โ Type
True error (0-1 loss, realizable case): D-probability of disagreement. This is what PACLearnable's success event measures.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ ENNReal
True error in โ: for use in bounds involving subtraction/absolute value. COUNTER-1 of TrueError. The toReal bridge loses information when the measure is โค.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ โ
VC dimension of a concept class: the size of the largest shattered set. Returns โโ = WithTop โ.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
A concept class is well-behaved if the ghost gap event is null-measurable. This is the minimal regularity assumption for the symmetrization proof.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Convert Bool labels to ยฑ1 reals. true โฆ 1, false โฆ -1.
Bool โ โ
The set of labelling patterns that C realizes on a finite sample S.
{X : Type u} โ ConceptClass X Bool โ (S : Finset X) โ Set (โฅS โ Bool)Uniform probability measure on a Fintype: (1/|X|) ยท count. This gives each point probability 1/|X|. Requires |X| > 0 (nonempty).
(X : Type u) โ [inst : MeasurableSpace X] โ [Fintype X] โ Nonempty X โ MeasureTheory.Measure X
The 0-1 loss for classification.
(Y : Type v) โ [DecidableEq Y] โ LossFunction Y
Online learning is strictly stronger than PAC learning.
(โ (X : Type) [inst : MeasurableSpace X] (C : ConceptClass X Bool) [MeasurableConceptClass X C],
OnlineLearnable X Bool C โ PACLearnable X C) โง
โ X x C, PACLearnable X C โง ยฌOnlineLearnable X Bool C(โ (X : Type) [inst : MeasurableSpace X] (C : ConceptClass X Bool) [MeasurableConceptClass X C],
OnlineLearnable X Bool C โ PACLearnable X C) โง
โ X x C, PACLearnable X C โง ยฌOnlineLearnable X Bool Cโ X x C, PACLearnable X C โง ยฌOnlineLearnable X Bool C
VCDim of threshold class on โ is finite (โค 1).
VCDim โ thresholdClassโ < โค
No 2-element subset of โ is shattered by the threshold class. Key: the labeling (smaller โ false, larger โ true) is impossible by monotonicity.
โ {S : Finset โ}, 2 โค S.card โ ยฌShatters โ thresholdClassโ SLittlestoneDim of threshold class = โค.
LittlestoneDim โ thresholdClassโ = โค
โ (d : โ), โ T, LTree.isShattered thresholdClassโ T
The threshold tree is shattered by any concept class containing all thresholds with indices in [lo, lo + 2^d - 1].
โ (lo d : โ) (C : ConceptClass โ Bool), (โ (n : โ), lo โค n โ n < lo + 2 ^ d โ (fun x => decide (x โค n)) โ C) โ LTree.isShattered C (thresholdTreeโ lo d)
Forward direction: OnlineLearnable โ LittlestoneDim < โค
โ (X : Type) (C : ConceptClass X Bool), OnlineLearnable X Bool C โ LittlestoneDim X C < โค
Relate mistakesFrom to the original mistakes function.
โ {X : Type} (L : OnlineLearner X Bool) (c : X โ Bool) (seq : List X), L.mistakesFrom L.init c seq = L.mistakes c seqCore adversary lemma.
โ {X : Type} (L : OnlineLearner X Bool) (s : L.State) {C : ConceptClass X Bool} {n : โ} (T : LTree X n),
LTree.isShattered C T โ Set.Nonempty C โ โ seq, โ c โ C, L.mistakesFrom s c seq = nHelper: shattering implies the concept class is nonempty.
โ {X : Type} {C : ConceptClass X Bool} {n : โ} (T : LTree X n), LTree.isShattered C T โ Set.Nonempty COnline learnable โ PAC learnable. ฮโโ: requires LittlestoneDim โ VCDim bridge or online-to-batch conversion.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool), OnlineLearnable X Bool C โ โ [MeasurableConceptClass X C], PACLearnable X C
Mistake-bounded learner โ VCDim โค M (universe-polymorphic).
โ {X : Type u} {C : ConceptClass X Bool} {M : โ}, MistakeBounded X Bool C M โ VCDim X C โค โMRelate mistakesFromU to the original mistakes function.
โ {X : Type u} (L : OnlineLearner X Bool) (c : X โ Bool) (seq : List X),
mistakesFromUโ L L.init c seq = L.mistakes c seqAdversary argument directly from shattering (universe-polymorphic). Given a shattered set S and any online learner L starting from state s, there exists a sequence and target concept where L makes |S| mistakes.
โ {X : Type u} (L : OnlineLearner X Bool) (s : L.State) {C : ConceptClass X Bool} {S : Finset X},
Shatters X C S โ โ seq, โ c โ C, mistakesFromUโ L s c seq = S.cardRestricted shattering: if C shatters S and we restrict to {c โ C | c x = b}, then S \ {x} is shattered by the restricted class (when x โ S).
โ {X : Type u} {C : ConceptClass X Bool} {S : Finset X},
Shatters X C S โ โ {x : X}, x โ S โ โ (b : Bool), Shatters X {c | c โ C โง c x = b} (S.erase x)VCDim < โค โ PACLearnable via UC route.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool),
VCDim X C < โค โ
(โ h โ C, Measurable h) โ (โ (c : Concept X Bool), Measurable c) โ WellBehavedVC X C โ PACLearnable X CFinite VCDim implies uniform convergence. Proof: VCDim < โ โ UC.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool),
VCDim X C < โค โ
(โ h โ C, Measurable h) โ (โ (c : Concept X Bool), Measurable c) โ WellBehavedVC X C โ HasUniformConvergence X CVCDim < โค โ growth function polynomially bounded by partial binomial sum. Forward direction of fundamental_theorem conjunct 5. Uses Sauer-Shelah: GrowthFunction(m) โค โ_{iโคd} C(m,i) where d = VCDim.
โ (X : Type u) (C : ConceptClass X Bool), VCDim X C < โค โ โ d, โ (m : โ), d โค m โ GrowthFunction X C m โค โ i โ Finset.range (d + 1), m.choose i
UC bad-event bound: for m โฅ mโ(v,ฮต,ฮด), the probability of the bad event (โ h with |TrueErr-EmpErr| โฅ ฮต) is at most ฮด. Composes symmetrization_uc_bound with growth_exp_le_delta.
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
ฮด < 1 โ
โ (v : โ),
0 < v โ
(โ (n : โ), v โค n โ GrowthFunction X C n โค โ i โ Finset.range (v + 1), n.choose i) โ
(16 * Real.exp 1 * (โv + 1) / ฮต ^ 2) ^ (v + 1) / ฮด โค โm โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
|TrueErrorReal X h c D -
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮต} โค
ENNReal.ofReal ฮดThe symmetrization uniform convergence bound: two-sided version. P[โhโC: |TrueErr-EmpErr| โฅ ฮต] โค 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8).
Proof strategy (4 steps):
1. Decompose absolute value: |TrueErr - EmpErr| โฅ ฮต โ (TrueErr - EmpErr โฅ ฮต) โจ (EmpErr - TrueErr โฅ ฮต)
have abs_decomp : โ (a b : โ),
|a - b| โฅ ฮต โ a - b โฅ ฮต โจ b - a โฅ ฮต := by
intro a b; constructor
ยท intro h; by_cases h' : a - b โฅ ฮต
ยท exact Or.inl h'
ยท exact Or.inr (by linarith [abs_sub_comm a b, le_abs_self (a - b)])
ยท intro h; cases h with
| inl h => exact le_trans (le_of_eq (abs_of_nonneg (by linarith))) (by linarith)
| inr h => exact le_trans (le_of_eq (abs_of_nonpos (by linarith) โธ ...)) ...2. Upper tail: P[โhโC: TrueErr-EmpErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step + double_sample_pattern_bound.3. Lower tail: P[โhโC: EmpErr-TrueErr โฅ ฮต] โค 2ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
symmetrization_step to the event EmpErr-TrueErr โฅ ฮต and bound the double-sample event {EmpErr_S - EmpErr_{S'} โฅ ฮต/2}. have swap_symmetry : DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S) - EmpErr(S') โฅ ฮต/2} = DoubleSampleMeasure D m {p | โ h โ C, EmpErr(S') - EmpErr(S) โฅ ฮต/2} := Measure.prod_swap ... 4. Union bound: P[|gap| โฅ ฮต] โค P[gap โฅ ฮต] + P[gap โค -ฮต] โค 2ยทGFยทexp(...) + 2ยทGFยทexp(...) = 4ยทGF(C,2m)ยทexp(-mฮตยฒ/8)
-- Uses: MeasureTheory.measure_union_le for the union of two events
-- CAST: 2 * X + 2 * X = 4 * X in ENNReal (need ENNReal.add_mul or similar)References: SSBD Theorem 6.7, Kakade-Tewari Lecture 19
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
|TrueErrorReal X h c D -
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮต} โค
ENNReal.ofReal (4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))Symmetrization step for the lower tail: P[โh: EmpErr-TrueErr โฅ ฮต] โค 2ยทP_{double}[โh: EmpErr_S-EmpErr_{S'} โฅ ฮต/2].
Mirror of symmetrization_step for the opposite direction. Uses hoeffding_one_sided_upper instead of hoeffding_one_sided.
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) - TrueErrorReal X h c D โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}Symmetrization: the probability of a large gap TrueErr-EmpErr is at most twice the probability of a large gap EmpErr'-EmpErr on the double sample.
Proof strategy (6 steps):
1. Witness selection: For S in the bad event, โh* โ C with TrueErr(h) - EmpErr_S(h) โฅ ฮต.
-- In the bad event set, extract h* by classical choice
have h_witness : โ xs โ bad_event, โ h* โ C,
TrueErrorReal X h* c D - EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โฅ ฮต2. Ghost sample mean: E_{S'}[EmpErr_{S'}(h)] = TrueErr(h) โฅ EmpErr_S(h*) + ฮต.
MeasureTheory.integral_pi to compute E[EmpErr] over product measure. have expected_emp_err : โ h* : Concept X Bool, โซ xs, EmpiricalError X Bool h* (sample xs) (zeroOneLoss Bool) โ(Measure.pi (fun _ : Fin m => D)) = TrueErrorReal X h* c D := by ... 3. Hoeffding on ghost sample: P_{S'}[EmpErr_{S'}(h) < TrueErr(h) - ฮต/2] โค exp(-mฮตยฒ/2).
hoeffding_one_sided with t = ฮต/2.hm_large hypothesis ensures exp(-mฮตยฒ/2) < 1/2: 2ยทln2 โค mฮตยฒ โน mฮตยฒ/2 โฅ ln2 โน exp(-mฮตยฒ/2) โค 1/2. have hoeffding_ghost : โ h* โ C, Measure.pi (fun _ : Fin m => D) {xs' | EmpiricalError X Bool h* (sample xs') (zeroOneLoss Bool) < TrueErrorReal X h* c D - ฮต/2} โค ENNReal.ofReal (Real.exp (-m * (ฮต/2)^2 * 2)) := by intro h* _; exact hoeffding_one_sided D h* c m hm (ฮต/2) (by linarith) (by ...) (by ...) 4. Complementary probability: P_{S'}[EmpErr_{S'}(h) - EmpErr_S(h) โฅ ฮต/2] โฅ 1/2.
5. Conditional to unconditional: The witness h* from step 1 also witnesses the double-sample event โhโC: EmpErr'-EmpErr โฅ ฮต/2. So: P_{S'}[double event | S bad] โฅ 1/2.
have conditional_bound : โ xs โ bad_event,
Measure.pi (fun _ : Fin m => D)
{xs' | โ h โ C, EmpiricalError ... xs' - EmpiricalError ... xs โฅ ฮต/2}
โฅ ENNReal.ofReal (1/2) := by ...6. Fubini integration: By Measure.prod_apply and Fubini: P_{S,S'}[double event] = โซ_S P_{S'}[double event | S] โฅ (1/2) ยท P_S[bad event] โน P_S[bad event] โค 2 ยท P_{S,S'}[double event].
-- Uses: MeasureTheory.Measure.prod_apply or lintegral_prod
-- MEASURABILITY: the double-sample event is measurable as a finite union
-- of sets of the form {(xs,xs') | EmpErr'(h) - EmpErr(h) โฅ ฮต/2} for h โ C.
-- Since C may be infinite, measurability requires care: the sup over h
-- must be shown to be measurable. For finite restriction patterns (โค 2^m
-- on Fin m โ Bool), this is a finite union.MEASURABILITY CONCERNS:
{xs | โ h โ C, ...} is NOT obviously measurable for infinite C. Strategy: decompose via restriction patterns. On any fixed xs, the set of labelings {(h(xs 0), ..., h(xs(m-1))) | h โ C} has at most GF(C,m) โค 2^m elements. So the โh event is a finite union of measurable sets.EmpiricalError is a finite sum of measurable functions, hence measurable.References: SSBD Lemma 4.5, Kakade-Tewari Lecture 19 Lemma 1
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
2 * Real.log 2 โค โm * ฮต ^ 2 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ h โ C,
TrueErrorReal X h c D - EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ
ฮต} โค
2 *
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}On the double sample, the probability that any hypothesis has EmpErr' - EmpErr โฅ ฮต/2 is bounded by GF(C,2m) ยท exp(-mฮตยฒ/8).
Proof strategy (Approach A โ standard exchangeability, 5 steps):
1. EXCHANGEABILITY: Under D^m โ D^m, the 2m draws zโ,...,z_{2m} are iid from D. The joint distribution is invariant under permutations of {1,...,2m}.
Key lemma: P_{D^mโD^m}[event(S,S')] = E_z[P_{split}[event | z]] where z = merged sample and the split is uniformly random among all C(2m,m) ways to partition z into two groups of m.
-- Measure.pi permutation invariance
have pi_perm_invariant : โ (ฯ : Equiv.Perm (Fin (2*m))),
(Measure.pi (fun _ : Fin (2*m) => D)).map (fun z i => z (ฯ i))
= Measure.pi (fun _ : Fin (2*m) => D) := by ...
-- Consequence: the event probability equals the split-averaged probability
have exchangeability :
DoubleSampleMeasure D m {p | โ h โ C, gap(p) โฅ ฮต/2}
= โซ z, SplitMeasure m {vs | โ h โ C, gap(split z vs) โฅ ฮต/2}
โ(Measure.pi (fun _ : Fin (2*m) => D)) := by ...2. CONDITIONING: For fixed merged sample z of 2m points:
-- Number of distinct patterns
have num_patterns : โ (z : MergedSample X m),
Set.ncard {p : Fin (2*m) โ Bool | โ h โ C, โ i, p i = (h (z i) โ c (z i))}
โค GrowthFunction X C (2*m) := by ...3. PER-PATTERN HOEFFDING ON SPLITS: For fixed z and fixed pattern p: Under uniformly random split (S,S') of z into two groups of m: diff(p, split) = (1/m) โ_{iโS'} a_i - (1/m) โ_{iโS} a_i
This is a function of the random partition. By Hoeffding's inequality for sampling without replacement (Serfling 1974): P_split[diff โฅ ฮต/2] โค exp(-mฮตยฒ/8)
Alternative derivation: Hoeffding without replacement from Hoeffding with replacement (iid signs) via coupling. The without-replacement bound is actually TIGHTER (variance reduction), but the with-replacement bound suffices.
-- Per-pattern concentration
have per_pattern_bound : โ (z : MergedSample X m) (a : Fin (2*m) โ โ)
(ha : โ i, a i โ Set.Icc 0 1),
SplitMeasure m {vs | (1/m) * โ i โ second_group vs, a i
- (1/m) * โ i โ first_group vs, a i โฅ ฮต/2}
โค ENNReal.ofReal (Real.exp (-(m : โ) * (ฮต/2)^2 / 2)) := by ...
-- Note: m*(ฮต/2)^2/2 = mฮตยฒ/84. UNION BOUND: P_split[โ pattern: diff โฅ ฮต/2 | z] โค (number of patterns) ยท max_pattern P_split[diff โฅ ฮต/2] โค GF(C,2m) ยท exp(-mฮตยฒ/8)
have union_bound : โ (z : MergedSample X m),
SplitMeasure m {vs | โ h โ C, gap(split z vs, h) โฅ ฮต/2}
โค ENNReal.ofReal (GrowthFunction X C (2*m) * Real.exp (-(m : โ) * ฮต^2 / 8))
:= by ...5. INTEGRATE: P_{D^mโD^m}[event] = E_z[P_split[event|z]] (by step 1) โค E_z[GF(C,2m) ยท exp(-mฮตยฒ/8)] (by step 4, pointwise) = GF(C,2m) ยท exp(-mฮตยฒ/8) (bound is independent of z)
-- The bound is a constant, so integrating gives the same constant
-- (using IsProbabilityMeasure for the 2m-fold product)Infrastructure needed:
Fin.sumFinEquiv : Fin m โ Fin n โ Fin (m + n) (available in Mathlib)mergeSamples / splitMergedSample (defined above)SplitMeasure and ValidSplit (defined above)Measure.pi permutation invariance (to be proved or imported)GrowthFunction on 2m points + sauer_shelah_exp_bound from Rademacher.leanMEASURABILITY CONCERNS:
GrowthFunction X C (2*m) is a natural number (deterministic), no measurability issue.References: SSBD Theorem 6.7, Hoeffding (1963), Serfling (1974)
โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D))
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ {X : Type u} [inst : MeasurableSpace X] [Infinite X] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (C : ConceptClass X Bool) (c : Concept X Bool),
(โ h โ C, Measurable h) โ
Measurable c โ
โ (m : โ),
0 < m โ
โ (ฮต : โ),
0 < ฮต โ
ฮต โค 2 โ
Set.Nonempty C โ
MeasureTheory.NullMeasurableSet
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2}
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) โ
have ฮผ := MeasureTheory.Measure.pi fun x => D;
(ฮผ.prod ฮผ)
{p |
โ h โ C,
EmpiricalError X Bool h (fun i => (p.2 i, c (p.2 i))) (zeroOneLoss Bool) -
EmpiricalError X Bool h (fun i => (p.1 i, c (p.1 i))) (zeroOneLoss Bool) โฅ
ฮต / 2} โค
ENNReal.ofReal (โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)))โ (Y : Type v) {inst : DecidableEq Y} [inst_1 : DecidableEq Y] (a a_1 : Y),
a = a_1 โ โ (a_2 a_3 : Y), a_2 = a_3 โ zeroOneLoss Y a a_2 = zeroOneLoss Y a_1 a_3The number of distinct restriction patterns of C on any n points is at most GF(C,n). For z : Fin n โ X, define patterns(z) = {p : Fin n โ Bool | โ h โ C, โ i, p i = (h(z i) โ c(z i))}. Then patterns(z).ncard โค GrowthFunction X C n by definition of GrowthFunction.
โ {X : Type u} [MeasurableSpace X] [Infinite X] (C : ConceptClass X Bool) (c : Concept X Bool) (n : โ) (z : Fin n โ X),
{p | โ h โ C, โ (i : Fin n), p i = decide (h (z i) โ c (z i))}.ncard โค GrowthFunction X C nRademacher MGF bound.
โ {m : โ},
0 < m โ
โ (a : Fin m โ โ) (c : โ),
0 โค c โ
(โ (i : Fin m), |a i| โค c) โ
โ (t : โ),
0 โค t โ
1 / โ(Fintype.card (SignVector m)) * โ ฯ, Real.exp (t * (1 / โm * โ i, a i * boolToSign (ฯ i))) โค
Real.exp (t ^ 2 * c ^ 2 / (2 * โm))cosh(x) โค exp(xยฒ/2). Standard sub-Gaussian bound.
โ (x : โ), Real.cosh x โค Real.exp (x ^ 2 / 2)
Generic finite exchangeability bound. Given a measure-preserving family of transformations on a probability space, a NullMeasurableSet S, and a pointwise bound on the sum of preimage indicators, conclude ฮฝ(S) โค B.
โ {ฮฉ : Type u_1} {G : Type u_2} [inst : MeasurableSpace ฮฉ] [inst_1 : Fintype G] [Nonempty G]
{ฮฝ : MeasureTheory.Measure ฮฉ} [MeasureTheory.IsProbabilityMeasure ฮฝ] (T : G โ ฮฉ โ ฮฉ) (S : Set ฮฉ),
(โ (g : G), MeasureTheory.MeasurePreserving (T g) ฮฝ ฮฝ) โ
MeasureTheory.NullMeasurableSet S ฮฝ โ
โ (B : ENNReal), (โ (z : ฮฉ), โ g, (T g โปยน' S).indicator 1 z โค B * โ(Fintype.card G)) โ ฮฝ S โค Bโ {X : Type u} [MeasurableSpace X] (C : ConceptClass X Bool) (v : โ),
0 < v โ
โ (m : โ),
0 < m โ
โ (ฮต ฮด : โ),
0 < ฮต โ
0 < ฮด โ
ฮด < 1 โ
(โ (n : โ), v โค n โ GrowthFunction X C n โค โ i โ Finset.range (v + 1), n.choose i) โ
(16 * Real.exp 1 * (โv + 1) / ฮต ^ 2) ^ (v + 1) / ฮด โค โm โ
4 * โ(GrowthFunction X C (2 * m)) * Real.exp (-(โm * ฮต ^ 2 / 8)) โค ฮด โง 2 * Real.log 2 โค โm * ฮต ^ 2Pure combinatorial inequality: โ_{i=0}^d C(m,i) โค (em/d)^d for d โค m, d โฅ 1.
โ (d m : โ), 0 < d โ d โค m โ โ i โ Finset.range (d + 1), โ(m.choose i) โค (Real.exp 1 * โm / โd) ^ d
Key arithmetic lemma for PAC bound: for t > 0, t^d * exp(-t) โค (d+1)!/t. Follows from exp(t) โฅ t^(d+1)/(d+1)! (partial sum of Taylor series).
โ {d : โ} {t : โ}, 0 < t โ t ^ d * Real.exp (-t) โค โ(d + 1).factorial / tTrivial bound: GrowthFunction โค 2^n for all concept classes. Each restriction to an n-element set yields a function in S โ Bool, and there are at most 2^n such functions.
โ {X : Type u} (C : ConceptClass X Bool) (n : โ), GrowthFunction X C n โค 2 ^ nUpper-tail Hoeffding: for iid Bernoulli(p) draws, the empirical average overshoots the mean by โฅ t with probability โค exp(-2mtยฒ).
This is the mirror of hoeffding_one_sided (which bounds the lower tail). The proof uses the same sub-Gaussian machinery with Z_i = indicator(x_i) - p (instead of p - indicator(x_i)).
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ TrueErrorReal X h c D + t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))One-sided Hoeffding: for iid Bernoulli(p) draws, the empirical average undershoots the mean by โฅ t with probability โค exp(-2mtยฒ).
Proof strategy (3 steps):
1. MGF bound (Hoeffding's lemma): For X โ [0,1] with E[X] = p, E[exp(s(X-p))] โค exp(sยฒ/8).
cosh_le_exp_sq_half infrastructure in Rademacher.lean. have mgf_bound : โ (s : โ), โซ x, Real.exp (s * (indicator x - p)) โD โค Real.exp (s^2 / 8) := by ... 2. Product independence: E[exp(sยทโ(X_i-p))] = โ E[exp(s(X_i-p))] โค exp(msยฒ/8).
MeasureTheory.Measure.pi independence structure.Measure.pi integral factorization for product of functions.fun xs => Real.exp (s * โ i, f (xs i)) is measurable (composition of measurable functions). have product_bound : โ (s : โ), โซ xs, Real.exp (s * โ i, (indicator (xs i) - p)) โMeasure.pi (fun _ => D) โค Real.exp (m * s^2 / 8) := by ... 3. Exponential Markov + optimize: P[โ(X_i-p) โค -mt] = P[exp(-sยทโ(X_i-p)) โฅ exp(smt)] โค exp(-smt + msยฒ/8). Optimize over s: set s = 4t to get โค exp(-2mtยฒ).
have markov_step : โ (s : โ) (hs : 0 < s), Measure.pi (fun _ => D) {xs | โ i, (indicator (xs i) - p) โค -(m : โ) * t} โค ENNReal.ofReal (Real.exp (-(s * m * t) + m * s^2 / 8)) := by ... have optimize : Real.exp (-(4*t * m * t) + m * (4*t)^2 / 8) = Real.exp (-2 * m * t^2) := by ring_nf CAST ISSUES to watch:
m : โ needs cast to โ in the exponent: (m : โ)EmpiricalError returns โ, TrueErrorReal returns โ, good โ no ENNReal gapENNReal, the bound exp(-2mtยฒ) is โโฅ0โ via ENNReal.ofRealReferences: SSBD Lemma B.3, Hoeffding (1963)
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โค TrueErrorReal X h c D - t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))Uniform convergence implies PAC learnability via ERM. The ERM learner (which exists by ermLearn) achieves PAC learning when uniform convergence holds. This is the second half of vcdim_finite_imp_pac.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool), Set.Nonempty C โ HasUniformConvergence X C โ PACLearnable X C
Output is in the hypothesis space
โ {X : Type u} {Y : Type v} (self : BatchLearner X Y) {m : โ} (S : Fin m โ X ร Y), self.learn S โ self.hypothesesโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C],
โ c โ C, Measurable cEvery concept in C is measurable
โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableConceptClass X C],
โ h โ C, Measurable hโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C]
(c : Concept X Bool), Measurable cAll concepts X โ Bool are measurable (for disagreement sets)
โ {X : Type u} {inst : MeasurableSpace X} (C : ConceptClass X Bool) [self : MeasurableConceptClass X C]
(c : Concept X Bool), Measurable cโ {X : Type u} [inst : MeasurableSpace X] (C : ConceptClass X Bool) [h : MeasurableConceptClass X C], WellBehavedVC X CUniform convergence bad event is NullMeasurableSet
โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableConceptClass X C],
WellBehavedVC X CA batch learner (PAC paradigm): takes a finite sample, returns a hypothesis.
Type u โ Type v โ Type (max u v)
The learner's hypothesis space
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ HypothesisSpace X YThe learning algorithm: given a sample, produce a hypothesis
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ {m : โ} โ (Fin m โ X ร Y) โ Concept X YA concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
Empirical error: average loss on a finite sample.
(X : Type u) โ (Y : Type v) โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ LossFunction Y โ โA concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
Growth function (shattering coefficient): ฯ_C(m) = max_{|S|=m} |{c|_S : c โ C}|. For each m-element set S, counts the number of distinct restrictions of C to S, then takes the supremum over all such S.
(X : Type u) โ ConceptClass X Bool โ โ โ โ
Uniform convergence of empirical error to true error over a hypothesis class. This is the property that makes finite VCDim โ PAC learnability work. BPโ connects here: this is ONE of the five characterizations.
M-DefinitionRepair (ฮโโ โ ฮโโ): The mโ must be INDEPENDENT of D and c. That's what "uniform" means โ convergence is uniform over all distributions and all target concepts. The original definition had mโ depending on D and c, making uc_imp_pac unprovable (PACLearnable's mf must be independent of D, c). Repaired: โ mโ is now BEFORE โ D, โ c. This STRENGTHENS the definition (A5-valid: adds content, doesn't simplify).
(X : Type u) โ [MeasurableSpace X] โ HypothesisSpace X Bool โ Prop
A hypothesis space is a set of candidate concepts that the learner searches over. When H = C (the realizable case), every concept in the target class is available. When H โ C or H โ C, we are in the improper/agnostic regime.
Structurally identical to ConceptClass but semantically distinct: ConceptClass is the ground truth collection; HypothesisSpace is what the learner has access to.
Type u โ Type v โ Type (max v u)
A hypothesis h is consistent with labeled sample S.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ PropA complete binary Littlestone tree of depth n.
Type โ โ โ Type
Path-wise shattering for complete trees. Path B: leaf case requires C.Nonempty (NAโโ).
{X : Type} โ {n : โ} โ ConceptClass X Bool โ LTree X n โ PropLittlestone dimension: the maximum depth of a complete shattered tree. Path B: returns WithBot (WithTop โ) so Ldim(โ ) = โฅ (NAโโ).
(X : Type) โ ConceptClass X Bool โ WithBot (WithTop โ)
A concept class with the measure-theoretic regularity needed for PAC theory.
Bundles three conditions: 1. Every concept in C is measurable 2. All concepts are measurable (needed for disagreement set measurability) 3. The UC bad event satisfies NullMeasurableSet (WellBehavedVC)
Condition 3 is the deep one: for uncountable C, the existential {โ h โ C, |TrueErr - EmpErr| โฅ ฮต} is NOT MeasurableSet in general. WellBehavedVC asserts it is NullMeasurableSet, which suffices for integration (lintegral_indicator_oneโ).
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Mistake-bounded learning: the learner makes at most M mistakes on ANY sequence. No distribution assumption. Characterized by Littlestone dimension.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ ConceptClass X Y โ โ โ Prop
Online learnable: there exists a finite mistake bound.
(X : Type u) โ (Y : Type v) โ [DecidableEq Y] โ ConceptClass X Y โ Prop
An online learner: receives instances one at a time, makes predictions sequentially.
Type u โ Type v โ Type (max (max 1 u) v)
Internal state type
{X : Type u} โ {Y : Type v} โ OnlineLearner X Y โ TypeInitial state
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.StateHelper: run an online learner on a sequence, counting mistakes.
{X : Type u} โ {Y : Type v} โ [DecidableEq Y] โ OnlineLearner X Y โ Concept X Y โ List X โ โ{X : Type u} โ {Y : Type v} โ [DecidableEq Y] โ (L : OnlineLearner X Y) โ Concept X Y โ L.State โ List X โ โ โ โCount mistakes starting from state s.
{X : Type} โ (L : OnlineLearner X Bool) โ L.State โ (X โ Bool) โ List X โ โPredict: given current state and new instance, output a prediction
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.State โ X โ YUpdate: given current state, instance, and revealed true label, update state
{X : Type u} โ {Y : Type v} โ (self : OnlineLearner X Y) โ self.State โ X โ Y โ self.StatePAC (Probably Approximately Correct) learning. The central definition of computational learning theory.
Sample space: Fin m โ X with i.i.d. product measure D^m. Labels: derived deterministically from target concept c (realizable case). Error: D-probability of disagreement between learner output and c.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
A set S โ X is shattered by concept class C if every labeling of S is realized by some concept in C.
(X : Type u) โ ConceptClass X Bool โ Finset X โ Prop
โ โ Type
True error (0-1 loss, realizable case): D-probability of disagreement. This is what PACLearnable's success event measures.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ ENNReal
True error in โ: for use in bounds involving subtraction/absolute value. COUNTER-1 of TrueError. The toReal bridge loses information when the measure is โค.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ โ
VC dimension of a concept class: the size of the largest shattered set. Returns โโ = WithTop โ.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
A concept class is well-behaved if the ghost gap event is null-measurable. This is the minimal regularity assumption for the symmetrization proof.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
Convert Bool labels to ยฑ1 reals. true โฆ 1, false โฆ -1.
Bool โ โ
Count mistakes starting from state s (universe-polymorphic version).
{X : Type u} โ (L : OnlineLearner X Bool) โ L.State โ (X โ Bool) โ List X โ โThreshold concept class on โ: { (ยท โค n) | n : โ }. VCDim = 1 (PAC-learnable), LittlestoneDim = โ (not online-learnable).
ConceptClass โ Bool
Build a shattered Littlestone tree of depth d for the threshold class restricted to thresholds in interval [lo, lo + 2^d - 1]. The concept class parameter C should contain all thresholds (ยท โค n) for lo โค n โค lo + 2^d - 1. We show the tree is shattered by C when C โ these thresholds.
โ โ (d : โ) โ LTree โ d
The 0-1 loss for classification.
(Y : Type v) โ [DecidableEq Y] โ LossFunction Y
Advice elimination (Ben-David & Dichterman 1998): If C is PAC-learnable with concept-dependent advice from a FINITE set A (with measurability regularity), then C is PAC-learnable without advice.
Proof strategy: run the advice-augmented learner with each a โ A on a training portion of the sample, producing |A| candidate hypotheses. Use a validation portion to select the candidate with lowest empirical error. Union bound over |A| advice values + Hoeffding on validation controls total failure probability. Sample complexity: O(m_orig(ฮต/2, ฮด/(2|A|)) + log(|A|/ฮด)/ฮตยฒ).
The [Fintype A] constraint is essential: for infinite A, the theorem is false (no finite union bound). [Nonempty A] ensures the advice space is inhabited.
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) [MeasurableHypotheses X C] (A : Type u_1) [inst_2 : Fintype A] [inst_3 : Nonempty A], PACLearnableWithAdviceRegular X C A โ PACLearnable X C
โ (X : Type u) [inst : MeasurableSpace X] (C : ConceptClass X Bool) [MeasurableHypotheses X C] (A : Type u_1) [inst_2 : Fintype A] [inst_3 : Nonempty A], PACLearnableWithAdviceRegular X C A โ PACLearnable X C
Split D^{mโ+mโ} into D^{mโ} ร D^{mโ} via splitUsedEquiv.
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(mโ mโ : โ) (Success : Set ((Fin mโ โ X) ร (Fin mโ โ X))),
MeasurableSet Success โ
(MeasureTheory.Measure.pi fun x => D) (โ(splitUsedEquivโ mโ mโ) โปยน' Success) =
((MeasureTheory.Measure.pi fun x => D).prod (MeasureTheory.Measure.pi fun x => D)) Successโ {X : Type u} [inst : MeasurableSpace X] {A : Type u_1} [inst_1 : Fintype A] [inst_2 : Nonempty A]
(cand : A โ Concept X Bool) (c : Concept X Bool) (D : MeasureTheory.Measure X) {m : โ} (Sval : Fin m โ X ร Bool)
(ฮท ฯ : โ),
0 โค ฮท โ
(โ (a : A), |TrueErrorReal X (cand a) c D - EmpiricalError X Bool (cand a) Sval (zeroOneLoss Bool)| โค ฮท) โ
โ (aStar : A),
TrueErrorReal X (cand aStar) c D โค ฯ โ TrueErrorReal X (cand (bestAdvice cand Sval)) c D โค ฯ + 2 * ฮทFor a probability measure, ฮผ(S) โฅ 1 - ฮผ(Sแถ), and hence ฮผ(S) โฅ 1 - ฮด if ฮผ(Sแถ) โค ฮด.
โ {ฮฉ : Type u_1} [inst : MeasurableSpace ฮฉ] (ฮผ : MeasureTheory.Measure ฮฉ) [MeasureTheory.IsProbabilityMeasure ฮผ]
(S : Set ฮฉ) (ฮด : ENNReal), ฮผ Sแถ โค ฮด โ ฮผ S โฅ 1 - ฮดSampling Nat.pair mโ mโ coordinates and taking the first mโ+mโ gives the same measure as sampling mโ+mโ coordinates directly. The extra junk coordinates integrate out via pi_cylinder_set_eq.
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
[MeasureTheory.SigmaFinite D] (mโ mโ : โ) (Success : Set (Fin (mโ + mโ) โ X)),
MeasurableSet Success โ
(MeasureTheory.Measure.pi fun x => D) (usedPrefixโ mโ mโ โปยน' Success) =
(MeasureTheory.Measure.pi fun x => D) SuccessCylinder set measure on product: if an event depends only on the first coordinates (those satisfying predicate p), then its measure under D^ฮน equals D^{p}(event). Uses piEquivPiSubtypeProd: D^ฮน โ D^{p} ร D^{ยฌp}, and (D^{p} ร D^{ยฌp})(A ร univ) = D^{p}(A) ยท D^{ยฌp}(univ) = D^{p}(A).
โ {ฮน : Type u_1} [inst : Fintype ฮน] [DecidableEq ฮน] {X : Type u} [inst_2 : MeasurableSpace X]
(D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D] [MeasureTheory.SigmaFinite D] (p : ฮน โ Prop)
[inst_5 : DecidablePred p] (S : Set ({ i // p i } โ X)),
MeasurableSet S โ
(MeasureTheory.Measure.pi fun x => D) {xs | (fun i => xs โi) โ S} = (MeasureTheory.Measure.pi fun x => D) Sโ {X : Type u} [inst : MeasurableSpace X] {A : Type u_1} (LA : LearnerWithAdvice X Bool A),
AdviceEvalMeasurable LA โ โ (a : A) {m : โ} (S : Fin m โ X ร Bool), Measurable (LA.learnWithAdvice a S)โ {X : Type u} [inst : MeasurableSpace X] {A : Type u_1} [inst_1 : Fintype A] (D : MeasureTheory.Measure X)
[MeasureTheory.IsProbabilityMeasure D] (c : Concept X Bool),
Measurable c โ
โ (cand : A โ Concept X Bool),
(โ (a : A), Measurable (cand a)) โ
โ (m : โ),
0 < m โ
โ (ฮท : โ),
0 < ฮท โ
ฮท โค 1 โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
โ a,
|TrueErrorReal X (cand a) c D -
EmpiricalError X Bool (cand a) (fun i => (xs i, c (xs i))) (zeroOneLoss Bool)| โฅ
ฮท} โค
ENNReal.ofReal (โ(Fintype.card A) * 2 * Real.exp (-2 * โm * ฮท ^ 2))Upper-tail Hoeffding: for iid Bernoulli(p) draws, the empirical average overshoots the mean by โฅ t with probability โค exp(-2mtยฒ).
This is the mirror of hoeffding_one_sided (which bounds the lower tail). The proof uses the same sub-Gaussian machinery with Z_i = indicator(x_i) - p (instead of p - indicator(x_i)).
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โฅ TrueErrorReal X h c D + t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))One-sided Hoeffding: for iid Bernoulli(p) draws, the empirical average undershoots the mean by โฅ t with probability โค exp(-2mtยฒ).
Proof strategy (3 steps):
1. MGF bound (Hoeffding's lemma): For X โ [0,1] with E[X] = p, E[exp(s(X-p))] โค exp(sยฒ/8).
cosh_le_exp_sq_half infrastructure in Rademacher.lean. have mgf_bound : โ (s : โ), โซ x, Real.exp (s * (indicator x - p)) โD โค Real.exp (s^2 / 8) := by ... 2. Product independence: E[exp(sยทโ(X_i-p))] = โ E[exp(s(X_i-p))] โค exp(msยฒ/8).
MeasureTheory.Measure.pi independence structure.Measure.pi integral factorization for product of functions.fun xs => Real.exp (s * โ i, f (xs i)) is measurable (composition of measurable functions). have product_bound : โ (s : โ), โซ xs, Real.exp (s * โ i, (indicator (xs i) - p)) โMeasure.pi (fun _ => D) โค Real.exp (m * s^2 / 8) := by ... 3. Exponential Markov + optimize: P[โ(X_i-p) โค -mt] = P[exp(-sยทโ(X_i-p)) โฅ exp(smt)] โค exp(-smt + msยฒ/8). Optimize over s: set s = 4t to get โค exp(-2mtยฒ).
have markov_step : โ (s : โ) (hs : 0 < s), Measure.pi (fun _ => D) {xs | โ i, (indicator (xs i) - p) โค -(m : โ) * t} โค ENNReal.ofReal (Real.exp (-(s * m * t) + m * s^2 / 8)) := by ... have optimize : Real.exp (-(4*t * m * t) + m * (4*t)^2 / 8) = Real.exp (-2 * m * t^2) := by ring_nf CAST ISSUES to watch:
m : โ needs cast to โ in the exponent: (m : โ)EmpiricalError returns โ, TrueErrorReal returns โ, good โ no ENNReal gapENNReal, the bound exp(-2mtยฒ) is โโฅ0โ via ENNReal.ofRealReferences: SSBD Lemma B.3, Hoeffding (1963)
โ {X : Type u} [inst : MeasurableSpace X] (D : MeasureTheory.Measure X) [MeasureTheory.IsProbabilityMeasure D]
(h c : Concept X Bool) (m : โ),
0 < m โ
โ (t : โ),
0 < t โ
t โค 1 โ
MeasurableSet {x | h x โ c x} โ
(MeasureTheory.Measure.pi fun x => D)
{xs |
EmpiricalError X Bool h (fun i => (xs i, c (xs i))) (zeroOneLoss Bool) โค TrueErrorReal X h c D - t} โค
ENNReal.ofReal (Real.exp (-2 * โm * t ^ 2))โ {X : Type u} [inst : MeasurableSpace X] {A : Type u_1} [inst_1 : Fintype A] [inst_2 : Nonempty A]
(cand cand_1 : A โ Concept X Bool),
cand = cand_1 โ
โ {m : โ} (Sval Sval_1 : Fin m โ X ร Bool), Sval = Sval_1 โ bestAdvice cand Sval = bestAdvice cand_1 Sval_1โ {X : Type u} {inst : MeasurableSpace X} {C : ConceptClass X Bool} [self : MeasurableHypotheses X C],
โ h โ C, Measurable hPACLearnableWithAdviceRegular X C A
Joint measurability of a sample-dependent advice learner's evaluation map.
{X : Type u} โ [MeasurableSpace X] โ {A : Type u_1} โ LearnerWithAdvice X Bool A โ PropA batch learner (PAC paradigm): takes a finite sample, returns a hypothesis.
Type u โ Type v โ Type (max u v)
The learning algorithm: given a sample, produce a hypothesis
{X : Type u} โ {Y : Type v} โ BatchLearner X Y โ {m : โ} โ (Fin m โ X ร Y) โ Concept X YA concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
Empirical error: average loss on a finite sample.
(X : Type u) โ (Y : Type v) โ Concept X Y โ {m : โ} โ (Fin m โ X ร Y) โ LossFunction Y โ โA concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
A learner augmented with advice.
Type u โ Type v โ Type u_1 โ Type (max (max u u_1) v)
Advice-augmented learning: advice โ sample โ hypothesis
{X : Type u} โ {Y : Type v} โ {A : Type u_1} โ LearnerWithAdvice X Y A โ A โ {m : โ} โ (Fin m โ X ร Y) โ Concept X YEvery concept in C is a measurable function. Krapp-Wirth precondition: ฮ(h) โ ฮฃ_Z for all h โ H.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
PAC (Probably Approximately Correct) learning. The central definition of computational learning theory.
Sample space: Fin m โ X with i.i.d. product measure D^m. Labels: derived deterministically from target concept c (realizable case). Error: D-probability of disagreement between learner output and c.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ Prop
PAC learnability with finite advice, plus measurability for holdout validation.
(X : Type u) โ [MeasurableSpace X] โ ConceptClass X Bool โ (A : Type u_1) โ [Fintype A] โ [Nonempty A] โ Prop
True error (0-1 loss, realizable case): D-probability of disagreement. This is what PACLearnable's success event measures.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ ENNReal
True error in โ: for use in bounds involving subtraction/absolute value. COUNTER-1 of TrueError. The toReal bridge loses information when the measure is โค.
(X : Type u) โ [inst : MeasurableSpace X] โ Concept X Bool โ Concept X Bool โ MeasureTheory.Measure X โ โ
Choose the advice value with minimum validation empirical error.
{X : Type u} โ
[MeasurableSpace X] โ
{A : Type u_1} โ [Fintype A] โ [Nonempty A] โ (A โ Concept X Bool) โ {m : โ} โ (Fin m โ X ร Bool) โ ASplit Fin (mโ + mโ) โ X into (Fin mโ โ X) ร (Fin mโ โ X) measurably.
{X : Type u} โ [inst : MeasurableSpace X] โ (mโ mโ : โ) โ (Fin (mโ + mโ) โ X) โแต (Fin mโ โ X) ร (Fin mโ โ X)Extract the first mโ + mโ coordinates from a sample of size Nat.pair mโ mโ.
{X : Type u} โ [MeasurableSpace X] โ (mโ mโ : โ) โ (Fin (Nat.pair mโ mโ) โ X) โ Fin (mโ + mโ) โ XThe 0-1 loss for classification.
(Y : Type v) โ [DecidableEq Y] โ LossFunction Y
VC dimension of homogeneous linear halfspaces is exactly n (Cover 1965; VapnikโChervonenkis). The class signClass (coordSpace n) of homogeneous linear halfspaces of โโฟ has VC dimension equal to the ambient dimension n: the Dudley bound gives โค n, and the n standard basis points are shattered, giving โฅ n.
โ (n : โ), VCDim (Fin n โ โ) (signClass (FLT.Halfspace.coordSpace n)) = โn
โ (n : โ), VCDim (Fin n โ โ) (signClass (FLT.Halfspace.coordSpace n)) = โn
VC dimension of homogeneous linear halfspaces: the Dudley upper bound.
The class of homogeneous linear halfspaces x โฆ (0 < โจw, xโฉ) in โโฟ has VC dimension at most n. This is the canonical instance of Dudley's bound (vcDim_signClass_le): the underlying function space is coordSpace n, of dimension n. (Cover 1965; VapnikโChervonenkis.)
โ (n : โ), VCDim (Fin n โ โ) (signClass (FLT.Halfspace.coordSpace n)) โค โn
Dudley's bound. The VC dimension of the linear sign class of a finite-dimensional subspace V โค (X โ โ) is at most dim V.
โ {X : Type u} (V : Submodule โ (X โ โ)) [FiniteDimensional โ โฅV], VCDim X (signClass V) โค โ(Module.finrank โ โฅV)The dimension of the coordinate space is n. The coordinate functionals form the dual basis of โโฟ, so their span has dimension n.
โ (n : โ), Module.finrank โ โฅ(FLT.Halfspace.coordSpace n) = n
The coordinate functionals are linearly independent. They are the dual basis of the standard basis; concretely, evaluating a vanishing combination at each standard basis point Pi.single j 1 isolates the j-th coefficient.
โ (n : โ), LinearIndependent โ (FLT.Halfspace.coord n)
Evaluating the i-th coordinate functional at the j-th standard basis point gives the identity matrix: coord n i (Pi.single j 1) = if i = j then 1 else 0.
โ (n : โ) (i j : Fin n), FLT.Halfspace.coord n i (Pi.single j 1) = if i = j then 1 else 0
The coordinate space is finite-dimensional (it is the span of a finite family).
โ (n : โ), FiniteDimensional โ โฅ(FLT.Halfspace.coordSpace n)
VC dimension of homogeneous linear halfspaces: the lower bound. The n standard basis points are shattered, so the VC dimension is at least n.
โ (n : โ), โn โค VCDim (Fin n โ โ) (signClass (FLT.Halfspace.coordSpace n))
The standard basis points are shattered by homogeneous halfspaces. Given any labelling, the ยฑ1-weighted functional g x = โ i, w i * x i with w i = ยฑ1 selected by the label realises it: g (Pi.single j 1) = w j, and 0 < w j iff the label is true.
โ (n : โ), Shatters (Fin n โ โ) (signClass (FLT.Halfspace.coordSpace n)) (FLT.Halfspace.basisPoints n)
A weighted functional evaluated at a standard basis point returns that point's weight: (โ i, w i * (Pi.single j 1) i) = w j.
โ (n : โ) (w : Fin n โ โ) (j : Fin n), โ i, w i * Pi.single j 1 i = w j
A ยฑ1-weighted coordinate functional belongs to the coordinate space. Concretely, for any weights w : Fin n โ โ, the functional x โฆ โ i, w i * x i is a member of coordSpace n.
โ (n : โ) (w : Fin n โ โ), (fun x => โ i, w i * x i) โ FLT.Halfspace.coordSpace n
Each standard basis point lies in basisPoints n.
โ (n : โ) (i : Fin n), Pi.single i 1 โ FLT.Halfspace.basisPoints n
There are exactly n standard basis points.
โ (n : โ), (FLT.Halfspace.basisPoints n).card = n
The map i โฆ Pi.single i 1 is injective (the basis points are distinct).
โ (n : โ), Function.Injective fun i => Pi.single i 1
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
A concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
The n standard basis points {Pi.single i 1 | i : Fin n} of โโฟ. These are the witnesses for the VC-dimension lower bound.
(n : โ) โ Finset (Fin n โ โ)
The i-th coordinate functional x โฆ x i on Fin n โ โ, viewed as an element of the function space (Fin n โ โ) โ โ.
(n : โ) โ Fin n โ (Fin n โ โ) โ โ
The coordinate space: the span of the n coordinate functionals (fun x => x i) inside (Fin n โ โ) โ โ. This is the (dual) space of homogeneous linear functionals on โโฟ; its sign class is exactly the family of homogeneous linear halfspaces.
(n : โ) โ Submodule โ ((Fin n โ โ) โ โ)
A set S โ X is shattered by concept class C if every labeling of S is realized by some concept in C.
(X : Type u) โ ConceptClass X Bool โ Finset X โ Prop
VC dimension of a concept class: the size of the largest shattered set. Returns โโ = WithTop โ.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
Evaluation at a point x : X as a linear functional on V: g โฆ (g : X โ โ) x.
{X : Type u} โ (V : Submodule โ (X โ โ)) โ X โ โฅV โโ[โ] โThe sign-pattern concept class of a subspace V โค (X โ โ): all concepts of the form x โฆ decide (0 < g x) for some g โ V.
{X : Type u} โ Submodule โ (X โ โ) โ ConceptClass X BoolAssouad's dual VC bound. If VCDim X C โค d, then VCDim(dualClass C) โค 2^(d+1) โ 1.
โ {X : Type u} {C : ConceptClass X Bool} {d : โ}, VCDim X C โค โd โ VCDim (โC) (dualClass C) โค โ(2 ^ (d + 1) - 1)โ {X : Type u} {C : ConceptClass X Bool} {d : โ}, VCDim X C โค โd โ VCDim (โC) (dualClass C) โค โ(2 ^ (d + 1) - 1)Assouad's coding lemma. If the dual class shatters a set S of concepts with 2^(d+1) โค S.card, then C shatters some set of d + 1 points.
Index 2^(d+1) of the shattered concepts by bitstrings; for each coordinate k, dual shattering realizes the labelling "is this codeword in the k-th bit slice", which supplies a point x k reading off exactly that bit. These d + 1 points are then shattered by C.
โ {X : Type u} {C : ConceptClass X Bool} {d : โ} (S : Finset โC),
Shatters (โC) (dualClass C) S โ 2 ^ (d + 1) โค S.card โ โ T, T.card = d + 1 โง Shatters X C TCodeword (cube b).val is in cubeBitSlice cube k iff b k = true.
โ {X : Type u} {C : ConceptClass X Bool} {n : โ} {S : Finset โC} (cube : (Fin n โ Bool) โช โฅS) (b : Fin n โ Bool)
(k : Fin n), โ(cube b) โ cubeBitSliceโ cube k โ b k = trueVCDim X C โค โd
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
A concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
A set S โ X is shattered by concept class C if every labeling of S is realized by some concept in C.
(X : Type u) โ ConceptClass X Bool โ Finset X โ Prop
VC dimension of a concept class: the size of the largest shattered set. Returns โโ = WithTop โ.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
The image of {b : Fin n โ Bool | b k = true} under a cube embedding: the codewords whose k-th bit is set.
{X : Type u} โ {C : ConceptClass X Bool} โ {n : โ} โ {S : Finset โC} โ ((Fin n โ Bool) โช โฅS) โ Fin n โ Finset โCEmbedding (Fin n โ Bool) โช โฅS when 2 ^ n โค S.card, via (Fin n โ Bool) โ Fin (2 ^ n) โช Fin S.card โ โฅS.
{X : Type u} โ {C : ConceptClass X Bool} โ {n : โ} โ (S : Finset โC) โ 2 ^ n โค S.card โ (Fin n โ Bool) โช โฅSThe dual concept class: each point x : X gives the evaluation concept c โฆ c x on the new domain โฅC. The dual class collects these over all points.
{X : Type u} โ (C : ConceptClass X Bool) โ ConceptClass (โC) BoolAssouad's lower bound. For a finite-VC class, โlogโ VCDimโ โค VCDim(dualClass): the exponential blow-up under dualization is necessary, not merely permitted. Together with the proven upper bound vcDim_dualClass_le (VCDim(dual) โค 2^(VCDim+1) โ 1) this sandwiches the dual VC dimension between โlogโ dโ and 2^(d+1) โ 1.
Stated with d := VCDim C extracted as a natural number (finite by hypothesis). The proof picks a shattered set T with 2^(logโ d) โค |T| (which exists because 2^(logโ d) โค d โค |T| for the supremal shattered set) and applies pow_le_vcDim_imp_le_vcDim_dualClass.
โ {X : Type u} {C : ConceptClass X Bool} {d : โ}, VCDim X C = โd โ 0 < d โ โ(Nat.log 2 d) โค VCDim (โC) (dualClass C)โ {X : Type u} {C : ConceptClass X Bool} {d : โ}, VCDim X C = โd โ 0 < d โ โ(Nat.log 2 d) โค VCDim (โC) (dualClass C)Assouad's lower construction. If C shatters a set T of at least 2^k points, then dualClass C shatters a set of k evaluation concepts, so k โค VCDim(dualClass C).
Bitstrings index a 2^k-subset of T; coordinate j is realized by a concept c_j โ C reading off the j-th bit, and the dual class shatters {c_j} because the evaluation point indexed by a bitstring ฯ realizes exactly the labelling ฯ of the c_j.
โ {X : Type u} {C : ConceptClass X Bool} {k : โ} (T : Finset X),
Shatters X C T โ 2 ^ k โค T.card โ โk โค VCDim (โC) (dualClass C)Each evaluation concept belongs to the dual class.
โ {X : Type u} (C : ConceptClass X Bool) (x : X), evalConcept C x โ dualClass CVCDim X C = โd
0 < d
A concept class is a set of concepts. Used by every paradigm, complexity measure, and criterion.
Primary definition: Set of functions. Used for PAC/agnostic PAC where concept classes are sets over which VC dimension, Rademacher complexity, covering numbers, etc. are measured. Alternative definitions below for contexts requiring decidability, enumerability, or measurability.
Type u โ Type v โ Type (max v u)
A concept is a function from domain to label. This is the atomic unit that concept classes collect and learners try to approximate.
Type u โ Type v โ Type (max u v)
A set S โ X is shattered by concept class C if every labeling of S is realized by some concept in C.
(X : Type u) โ ConceptClass X Bool โ Finset X โ Prop
VC dimension of a concept class: the size of the largest shattered set. Returns โโ = WithTop โ.
(X : Type u) โ ConceptClass X Bool โ WithTop โ
The dual concept class: each point x : X gives the evaluation concept c โฆ c x on the new domain โฅC. The dual class collects these over all points.
{X : Type u} โ (C : ConceptClass X Bool) โ ConceptClass (โC) BoolThe evaluation concept at a point x: the map c โฆ c x on โฅC.
{X : Type u} โ (C : ConceptClass X Bool) โ X โ โC โ Bool