All Notes
Definition. Representable Grammars representable-grammar
For any string , we can define a representable grammar which matches exactly the string and nothing else. The parse trees for a representable grammar are proofs that the string is exactly equal to :
Predecessors simplify later predecessors-simplify-later
A predecessor of is a top element of its strict downset: a strict morphism through which every strict morphism into factors uniquely. Equivalently, β the downset is representable.
The Yoneda lemma then collapses later to evaluation:
with becoming restriction along . The name is from the naturals: every strict map into factors through , so on β in the topos of trees β later is just the shift
and a LΓΆb step is a base value together with a rule producing the value at from the value at .
When this predecessor exists, we can give a simpler description of later, as in the topos of trees, but this may not be possible in all direct categories.
Algebras as a displayed category algebras-displayed
Fix an endofunctor . The -algebras form a displayed category over .
Over an object , a displayed object of is a structure map
Over , a displayed morphism from to is the proposition that is an algebra homomorphism:
The total category is the category of -algebras.
The co-EilenbergβMoore category as a displayed category co-eilenberg-moore-displayed
A comonad on is a monad on . Everything about its coalgebras is then inherited from the EilenbergβMoore construction, instantiated at the opposite category β nothing is defined twice.
Algebras of over are coalgebras of the underlying endofunctor, and the monad algebra laws, read in , are the comonad coalgebra laws β the unit and multiplication of , viewed in , are the counit and comultiplication :
The co-EilenbergβMoore category is the opposite of the total category:
Coalgebras as a displayed category coalgebras-displayed
Fix an endofunctor . Coalgebras require no new construction: a coalgebra is an algebra in the opposite category. Define
a displayed category over , where is acting on the opposite category.
Concretely, over an object a displayed object is a structure map
The category of coalgebras is the opposite of the total category:
The outer opposite returns morphisms to the direction of : a morphism is a map with
Definition. The comparison functor of an adjunction comparison-functor
An adjunction with and induces a monad on . Write for the counit of the adjunction. Every object of then induces a -algebra carried by the object , witnessed by the map
This assignment extends to a functor into the EilenbergβMoore category,
the comparison functor of the adjunction.
Dually, an adjunction induces a comonad on the other side and a comparison into the co-EilenbergβMoore category. When these comparisons are equivalences we say that the adjunction is (co)monadic.
Definition. Direct categories direct-category
A well-founded order is a set with a proposition-valued transitive relation admitting no infinite descent: every element is accessible. Write for .
A direct structure on a category over is a functor
into the well-founded order viewed as a poset category. The functor organizes two pieces of data at once: an ordering on the objects, and the invariant that morphisms respect it β forces .
A direct structure equips the objects with a well-founded strict relation
Intuitively, direct categories are the right generalization of well-foundedness to the categorical setting: a direct category is essentially one whose underlying graph is a directed acyclic graph, layered by degree, so that data at an object may be defined by recursion from data at all objects strictly below it.
The degrees order the objects, while the morphisms of say how an object sits over its predecessors.
Definition. Earlier on presheaves earlier-presheaf
Later takes a limit over smaller indices. Dually, the earlier modality takes a colimit over larger indices: an element of at is a -element sitting at some object strictly above , carried down along a chosen morphism .
Earlier is left adjoint to later:
Under this adjunction, corresponds to
.
The EilenbergβMoore category as a displayed category eilenberg-moore-displayed
Fix a monad on . Its EilenbergβMoore category arises in two displayed layers. The first layer is the displayed category of algebras of the underlying endofunctor.
The second layer, , is displayed over the total category . Over an algebra the displayed objects are the propositions that satisfies the monad algebra laws:
The EilenbergβMoore category is the total category of the tower:
Definition. Initial algebra initial-algebra
Fix an endofunctor . An initial -algebra, written , is an initial object of the category of algebras .
Unfolding the universal property: an initial algebra is an algebra such that every algebra admits a unique morphism satisfying
Definition. Later on families later-family
Conjugation with the adjunction between presheaves and families lets us induce a later construction on families from the one on presheaves,
Concretely, later on families evaluates to
Definition. Later on presheaves later-presheaf
The later modality for presheaves on a direct category is given by the presheaf of natural transformations
out of the strict downset.
An element of at is a coherent choice of -elements at all objects strictly smaller than .
Restriction in along precomposes with the induced map . At an object of minimal degree the strict downset is empty, so is trivial there.
Via functoriality, every presheaf restricts to smaller indices. Thus we may define the map
that sends an element over to the family of all its restrictions along morphisms from strictly lower objects.
Theorem. LΓΆb induction on families lob-family
Like later on families, the recursion principle for families is inherited from that on presheaves. Given a family and a step
the construction is a chain of transpositions with LΓΆb for presheaves used in the middle:
Just as for presheaves, the fixed point constructed above is unique: the two transpositions are bijections, and the presheaf-level fixed point is already unique.
Theorem. LΓΆb induction for presheaves on a direct category lob-presheaf
Let be a direct category and a presheaf on it. Every map
has a fixed point: a global element with
and this fixed point is unique.
The hypothesis says: the value of at any object is determined by its values over the strict past β turns a coherent family over the strict downset of into a value at . The proof is recursion along the well-founded : at each , the section already constructed over the past assembles into an element of , and extends it to .
Definition. Locally contractive endofunctors locally-contractive-functor
Write for the presheaf of morphisms . An endofunctor on presheaves is locally contractive when its action on morphisms factors through later: there is a map
Definition. Monadicity and comonadicity monadicity-comonadicity
An adjunction with and induces a monad on , and a comparison functor
sending each object of to the -algebra it carries.
The functor is monadic when is an equivalence: the adjunction exhibits as objects of equipped with algebraic structure for , the EilenbergβMoore category.
Comonadicity is monadicity in the opposite category: a left adjoint with right adjoint induces a comonad on , a comparison into the co-EilenbergβMoore category, and is comonadic when this comparison is an equivalence.
The adjoint triple between presheaves and families presheaf-family-adjoint-triple
A family over is a set for each object , with no action of morphisms. Families form a category : a morphism is a function for each .
Forgetting the restriction maps of a presheaf gives a functor
It has both a left and a right adjoint,
The two adjoints demonstrate different means of forcing a family to be functorial. The right adjoint universally quantifies over morphisms in,
with restriction along given by precomposition. The left adjoint instead existentially quantifiers over morphisms out:
with restriction acting on the first component. (For we ask that have a set of objects, so that this sum is a set and thus defines a presheaf.)
Theorem. Presheaves are monadic and comonadic over families presheaves-monadic-comonadic-over-families
The adjoint triple induces a monad and a comonad on .
Both comparison functors are equivalences: presheaves are the EilenbergβMoore algebras of and the co-EilenbergβMoore coalgebras of ,
So presheaves are both monadic and comonadic over families.
Reading the algebra structure concretely: a -algebra on a family is a map for each , subject to the monad algebra laws β that is, exactly a functorial action of restriction.
The comonadic reading is the same structure seen from the elementβs side: a -coalgebra is a map , giving each value its restriction along every morphism into . Where the monad says restriction acts on values, the comonad says a value already carries all of its restrictions β and the coalgebra laws say it does so coherently.
Definition. Proper and maximal sieves proper-maximal-sieve
The representable is itself a sieve on . A sieve on is proper when it is not equal to the representable.
Say that a proper sieve is maximal when it contains all other proper sieves as a sub-sieve.
Definition. Sieves sieve
A sieve on an object of is a subobject of the representable presheaf : a presheaf with a monic morphism . Sieves are a generalization from the notion of ideal found in ring theory to category theory.
A morphism belongs to , written , when lies in the image of the inclusion at . Because is a presheaf and the inclusion is natural, membership is closed under precomposition:
A sieve is thus a βdownward closedβ collection of morphisms into .
Sieves on are ordered by refinement: when every morphism belonging to belongs to .
Theorem. Maximality of the strict downset among proper sieves strict-downset-maximal
Call a direct structure reflecting when every morphism between objects of equal degree is a split epimorphism. In a reflecting direct category, every non-invertible-in-degree morphism strictly raises degree, and the strict downset is as large as a proper sieve can be:
If the direct structure is reflecting, then every proper sieve on refines into the strict downset:
Suppose with of equal degree. By reflection has a section , and closure under precomposition gives , contradicting properness.
So every morphism in strictly raises degree. That is, every morphism in is also a member of .
Definition. The strict downset sieve of a direct category strict-downset-sieve
Let carry a direct structure. The strict downset of an object is the presheaf of morphisms into from strictly lower objects:
with restriction by precomposition β well defined since degrees are non-decreasing, so precomposing can only stay strictly below.
The evident inclusion makes a sieve on . It is moreover a proper sieve, as it exlcudes the identity.
Definition. Terminal coalgebra terminal-coalgebra
Fix an endofunctor . A terminal -coalgebra, written , is a terminal object of the category of coalgebras β equivalently, an initial algebra for .
Unfolding the universal property: a terminal coalgebra is a coalgebra such that every coalgebra admits a unique morphism satisfying
Definition. Total category total-category
The total category of a displayed category over collects the displayed data into a single category. Its objects are pairs of an object of with an object over it, and its morphisms are pairs of a morphism with a displayed morphism over .
Projecting out the first components is a functor . Constructions presented displayed β algebras, EilenbergβMoore categories β get their forgetful functor for free as this projection.
Well-founded posets are thin direct categories well-founded-poset-as-thin
Every poset forms a thin category. Similarly, if the poset is well-founded then it induces a thin direct category.
Definition. Category category
A category consists of
- A type of objects
- For each pair of objects a set of morphisms . We may simply write a morphism with an arrow, denote as or or similar
- A composition operation on morphisms. For and , there is a morphism
- For each , an identity morphism
Left-unitality of composition: for all , an equality
Right-unitality of composition: for all , an equality
Associativity of composition: for all , , , an equality
Concretely, the definition above is meant to model the one used in the Cubical standard library [1].
However, the notion of category is flexible. Depending on the context, we may be talking of small, locally small, wild, or any other kind of category that may augment which things we require to be (homotopy) sets, which things we require to be small types, etc. For the most part, the same idea of a category will apply across all of these settings.
Definition. Displayed Category displayed-category
A displayed category over a base category packages the data of a category that βlies overβ : each object and each morphism of is equipped with a fiber of objects and morphisms displayed atop it. It consists of
- For each object , a type of displayed objects lying over . We write for a displayed object over .
- For each morphism and displayed objects and , a set of displayed morphisms lying over . We write such a displayed morphism as or , subscripting the arrow with the base morphism it lies over.
- A displayed composition operation. For and with displayed morphisms and , there is a displayed morphism lying over the composite .
- For each , a displayed identity lying over .
Left-unitality, displayed: for all , a heterogeneous equality
Right-unitality, displayed: for all , a heterogeneous equality
Associativity, displayed: for all , , over , , , a heterogeneous equality
Because the type of displayed morphisms depends on the base morphism , the two sides of each displayed law inhabit different displayed hom-sets β those indexed by the two sides of the corresponding base-category equation. Each displayed law is therefore a heterogeneous equality : a path from to lying over the base path , rather than an equation within a single fixed set.
Concretely, this definition models the one used in the Cubical standard library [1].
A displayed category is to a category as a dependent type is to a context. In this way, displayed categories are effectively dependent categories, as we present the structure of as a category parametrized by the structure of .
Displayed categories were introduced by Ahrens and Lumsdaine [2]. The point is that a displayed category over is equivalent to the data of a category together with a functor , but presented as families indexed by the objects and morphisms of β so that constructions like Grothendieck fibrations can be defined without ever invoking equality of objects.
Definition. Bicategory bicategory
A bicategory is a notion of weak 2-category that arises as a category weakly enriched in categories. That is, instead of having hom sets, between any two objects a bicategory has hom categories such that the enriched category laws hold up to invertible 2-cell rather than strictly.
A bicategory consists of
- A type of objects , or 0-cells
- For all , a category . We may elide the subscript and simply write this as . Refer to the objects of as 1-cells between and , and we may write as or . For , refer to the morphisms in between and as 2-cells and write the morphism as or
- For each , an identity 1-cell
- For all , a composition functor . For 1-cells and , write their composite as
For all , a natural isomorphism, the associator between the two composite functors that compose the leftmost, respectively rightmost, pair first:
Its component at 1-cells is the invertible 2-cell
For all , natural isomorphisms, the left unitor and right unitor , each between an endofunctor of and the identity functor:
where and in the pairings denote the constant functors at the identity 1-cells. The components at a 1-cell are the invertible 2-cells
such that for all and the triangle below commutes in :
and such that for all composable 1-cells the pentagon below commutes:
Definition. Monad in a bicategory monad-in-a-bicategory
Fix a bicategory , with composition , identity 1-cells , associator , and unitors . A monad in internalises the usual notion of monad: it is an endo-1-cell carrying a multiplication and a unit that satisfy the monoid laws up to the coherence cells of the bicategory.
A monad in consists of
- a 0-cell , the object the monad acts on;
- an endo-1-cell ;
- a multiplication 2-cell ;
- a unit 2-cell ;
such that is associative: the following diagram of 2-cells commutes in , where the top map is the associator that rebrackets the threefold composite:
and such that and satisfy the unit laws: the following two diagrams commute in , where the hypotenuses are the left and right unitors:
Taking to be the bicategory of categories, functors, and natural transformations recovers an ordinary monad on a category: is the endofunctor, the multiplication, and the unit, with the coherence cells all identities.
Freely transported terms in dependent type theory freely-transported-terms
Given and we can make sense of transported terms along equalities between indices in . Say, with
for .
For instance, if and then
To avoid landing in transport hell, I suspect that it may be preferable to work inside of a description of freely transported terms instead of taking semantic transports. The hypothesis is that by using descriptions of formal transport rather than actually computing a transport, we may defer the computation of an actual transport until the end of a construction. So instead of working with directly, perhaps we may work with
I think that this is very closely related to the Fording trick, as a map out of ,
can instead be described as a map,
Both this and the fording trick use the Coyoneda lemma to represent an dependent type family.
Definition. Displayed Total Category displayed-total-category
Given a displayed category over a category and another displayed category over , the total category of , we can define , the displayed total category of , as a displayed category over .
Definition. Free Monoidal Category over a Set free-monoidal-category
Fix a set . The objects of the free monoidal category over , , are generated inductively by the elements of and a unit element over a binary operation . The morphisms are given by a quotient-inductive type. They are generated by associators, unitors, and identity over composition and parallel action over then quotiented by associativity and composition equation to satisfy the category laws, equations constraining the associators/unitors to be natural isomorphisms, and pentagon/triangle equations to satiate the axioms of a monoidal category.
Definition 0.1. Global Elimination Principle for the Free Monoidal Category free-monoidal-category-elimination
Given any displayed monoidal category over with an interpretation , we may construct a global section . We refer to this as the global elimination principle of .
Definition. Global Elimination Principle for the Free Monoidal Category free-monoidal-category-elimination
Given any displayed monoidal category over with an interpretation , we may construct a global section . We refer to this as the global elimination principle of .
Products of Categories as Total Categories product-as-total-category
Given categories and , is equivalent to the total category of weakening.
Definition. Quiver quiver
A quiver is just a directed graph presented via a type of objects, a type of edges, and two projection functions that pick out source and target of an edge.
Definition. Reindexing a Displayed Category reindexing
Given a displayed category over a category and a functor , the reindexing of along is a displayed category over . It is simply the portion of over the image of .
Definition. Thin Category thin-category
A category is thin if there is at most one morphism between any two objects.
Definition. Weakening a Category weakening-category
For categories and , we can define the weakening of over as a displayed category over which trivially displays a copy of over each object of .
Definition. Element of a Presheaf element-of-presheaf
An element of a presheaf at an object is an element of the set .
Definition. Universal Element of a Presheaf universal-element
A universal element of a presheaf on a category is an element , where is some object of , demonstrating that is representable by .
(Slightly) more concretely, a universal element is captured by the following three pieces of data
- An object of
- An element
- A proof that the map sending a morphism to is an equivalence
This third point states that morphisms from into are uniquely determined by an element of at the domain .
Or equivalently, universal elements are terminal in the category of elements.
What is a universal property, really? universal-property
Universal properties are a convenient method for defining an object in a category up to isomorphism. Rather than giving a concrete, bottom-up construction of an object, we can instead uniquely specify its behavior.
Consider the example of products in a category . We say that the product of and is any object of such that the following diagram commutes.
We say that satisfies the universal property of the product of and .
Surely this matches our set-based intuition of what a product should behave like. Similarly, we can sketch out constructions of other universal properties like initial objects, terminal objects, exponentials, etc. However, what is precisely meant by the term universal property?
The notion of a universal property is made precise by the notion of a universal element of a presheaf. That is, an object satisfies a universal property if we can build a universal element of the appropriate presheaf at that object.
Letβs look at the universal element characterization of the products example. Note that a map into a product is determined by a map into each component. To map into , we need both a map into and a map into , as in the above diagram. That is, to build a map , we must simultaneously provide elements of and at .
Using the product of presheaves, this means we are providing a single element of the presheaf . Quite nicely, the universal element of this presheaf provides the object of that is the product of and . The universal element, provided that it exists, contains the following data:
- An object
- An element
- A proof that the map sending to a pair of maps , is an equivalence. Therefore, any element of factors through
Recall that the product of presheaves is computed pointwise in the category of sets, so if we expand the type of the element above we find that is a pair of maps and .
The first part of this pair is precisely . Correspondingly, the second part of this pair is . Finally, the universality of the element (i.e. the proof that any other element factors through ) captures our commutative diagram from above.
This is a very rough sketch of what a universal property is, and has elided for now an important application of the Yoneda lemma. In any case, all a universal property is really saying is that a particular presheaf is representable; and, rather elegantly, a universal element of a presheaf is convenient packaging of that representability proof.
In summary, universal properties are not as ad-hoc as they may initially seem, and the language of presheaves provides a reusable and precise definition that can be instantiated to describe a very large class of properties.
Cloudflare Parsing Error cloudflare-parsing-error
An error in the Cloudflare HTML parser would allow uninitialized memory to be dumped when there were imbalanced HTML tags.
This is indeed a case where a verified parser would have alleviated the issue.
Definition. GrΓΆbner Basis groebner-basis
Let be a ring and a polynomial ring over it. Suppose is an ideal. A GrΓΆbner basis for is a generating set of polynomials for the ideal that is minimal with respect to a given ordering on the monomials .
Given an ideal , a GrΓΆbner basis for may be found via Buchbergerβs algorithm. Intuitively, Buchbergerβs algorithm attempts to solve a system of polynomial equations by iterated polynomial division to eliminate variables. At any given point, there is a degree of freedom in which what variable will be eliminated by the next division. The algorithm attempts to eliminate variables with respect to the monomial ordering.
Buchbergerβs algorithm may be viewed simultaneously as a generalization of the Quine-McCluskey Boolean minimization algorithm and as a special case of the Knuth-Bendix algorithm.
The complexity of Buchbergerβs algorithm is a little unwieldy to estimate in general. However, just like SAT solvers there are enough optimizations to execute Buchberger reasonably fast in practice. For instance, it is fast enough to handle several hundreds of polynomials, each having hundreds of terms with very large coefficients.
There are some very fun applications of this approach, such as Solving Sudoku with Algebra. This idea has also been applied to inferring polynomial loop invariants.
GrΓΆbner Bases for Inferring Polynomial Loop Invariants groebner-loop-invariants
When the tools in I4: Incremental inference of inductive invariants for verification of distributed protocols and On Symmetry and Quantification: A New Approach to Verify Distributed Protocols search for an inductive invariant of a distributed system, the search procedure instantiates a series of small finite models and tries to infer from their truth tables a series of logical formulae that hold over those finite models. These formulae are found by running the Quine-McCluskey algorithm for minimization of Boolean functions. The prime implicants found by Quine-McCluskey have a latent symmetry that can be abstracted into quantified formulae. There are only so many small numbers, and so small finite models may propose formulae that do not hold at larger sizes. However, if you find a formula that holds at size as well as size , then it is likely a good candidate to hold at all sizes. You need to be a little careful if your protocol is indexed by several variables, but mostly this general idea holds when abstracting to a protocol of unbounded size. Further discussion of this idea can be found in SAT-based quantified symmetric minimization of the reachable states of distributed protocols: An update.
My observation was that Quine-McCluskey is just a special instance of Buchbergerβs algorithm for computing GrΓΆbner Bases. That is, you can describe Boolean formulae as polynomials over the field with two elements, and in this translation Quine-McCluskey and Buchberger each compute the same data. This observation isnβt new in and of itself, but it does open up an opportunity to generalize the invariant search procedure that is used above.
The place I went looking to apply this idea was in the search of polynomial loop invariants. If a loop had an invariant that is expressible as a polynomial relation between the program variables, then you could apply the same idea as above to infer the loop invariant.
Suppose are the variables in scope of program and contains a loop that we want to infer an invariant for. Denote the value of at the -th loop iteration by . The invariant search procedure proceeds intuitively as the following: we will keep track of the minimal set of polynomials that could interpolate between all of the variable assignments that we have witnessed thus far. Formally this is kept track of by the ideal of polynomials. At the -th loop iteration we add a new generator to the ideal which corresponds to the assignments . The GrΓΆbner basis for this ideal provides the minimal data needed to generate all the assignments witnessed thus far, so if this process saturates then the GrΓΆbner basis encodes a polynomial loop invariant. The nice thing about polynomials is that they have finite degree which guarantees that this process does indeed saturate (provided that the degree of the invariant is smaller than the number of loop iterations).
I was so excited to find this idea. Iβd felt like it was my first good idea in grad school. Then I read Automatic Generation of Polynomial Loop Invariants: Algebraic Foundations and found out someone had done this 20 years ago. I still wonder from time to time if there is room to further refine this idea or perhaps further generalize it. For instance, there is further generalization beyond Quine-McCluskey or Buchberger to the Knuth-Bendix algorithm, which seems to be a more general instance of both of these algorithms. So perhaps this search procedure can be weakened to an even more general class? Although, Iβm not yet familiar much with the Knuth-Bendix algorithm.
A Method for Verifying Translational Invariance of Image Processing Neural Networks translation-invariance-verification
My understanding for how one may prove a safety property for a neural network is as follows. First, express as an input-output property. That is, choose some region in the domain of to represent the inputs of interest. Further choose some region in the codomain of to describe safe outputs. Then describe as the property:
For instance, this is more or less how Reluplex, Marabou, and -CROWN each work. Squinting my eyes, this is the only such method for verifying a safety property for . That is, this is the only way to get a 100% guarantee that holds rather than some high measure of confidence.
I believe in this formalism I have an idea for how to express the property β is translationally invariantβ where is an image classifying net.
To this end, we need a continuous artifact that captures what it means to translate an image. We can express an image as a matrix of pixels (or perhaps several parallel matrices if we care about color channels, but stick to a single grayscale matrix for now). To shift over by a single pixel, we may left-multiply by the shift matrix , where is the matrix filled with zeros and has 1β²s on the subdiagonal. Note that all the directions of shifting are captured by the combinations of left/right multiplication by , .
Shifting by multiple pixels is now expressed by the matrix , however this is still a discrete dynamical system. We donβt have a continuous object by which we can test our safety property. My initial thought to continuousify this system was to express some sort of exponential flow by . That is, consider the matrix
then captures what it means to translate the image over by , a real-valued βamount of shiftingβ. This choice of continuous artifact could then be used for a verification effort, however there are some issues related to numerical stability because isnβt invertible. Because it is not invertible, doesnβt actually exist. So to make the above construction work, we need to mildly perturb and take a pseudoinverse. This sort of works and makes it so does capture some real valued shift, but it is only accurate when is small. So this is maybe problematic for our verification effort. This may be resolvable by chopping up the problem into subproblems, each of which is thin enough for the current iteration is accurate enough on that subproblem. However, there are two big things to consider.
- The exponential approach above is probably too complicated. It seems
likely that you may be able to take some (sequence of) linear interpolation(s) between the βs. This will still be some continuous object that captures a real-valued shift and it will be much more stable than the approach above.
- Even if we sort out which continuous object represents translation
of an image, I cannot for the life of me train any neural network that is translationally invariant. So the proof method is useless if there is nothing that it would ever apply to.
Precisely in this last point, I have mostly focused on trying to train a small CNN for MNIST handwritten digit classification that preserves the output class for small, reasonable translations of the digit. Iβve used data augmentation to predispose the network to being translationally invariant, and even though I can get a high degree of invariance, I cannot get a network that is invariant for all of the examples even in the training set or a reserved testing set.
It may be the case that the method I propose for measuring this invariance could be used to adversarially train a network to have better invariance. It may also be the case that all CNNs are bound to suffer from small degrees of translational sensitivity. I cannot find the citation at the moment, but there was a paper that suggested that CNNs suffer from weird issues of translational sensitivity that relate to the size of the convolutional window. So maybe this approach is doomed to fail anyway.
On the whole, I will say that machine learning verification almost sounds like an oxymoron. That is, if you have the expressivity to properly state a sophisticated safety property, then you likely understand the problem enough to not need to resort to machine learning in the first place. So almost tautologically, it seems that there cannot be satisfying verification of neural nets, as the tasks of machine learning and verification live on very different epistemic foundations.
The related works I could liberate from Zotero may be found below.
15 entries
0.2 Adequate Losses via Quantitative Linear Logic capucci-2026-adequate
0.3 Quantitative Linear Logic for Neuro-Symbolic Learning and Verification flinkow-2026-quantitative
0.4 Quantifiers for Differentiable Logics in Rocq (Extended Abstract) marulandagiraldo-2025-quantifiers
0.5 Fundamental Components of Deep Learning: A category-theoretic approach gavranovicFundamentalComponentsDeep
0.6 Architecture-Preserving Provable Repair of Deep Neural Networks taoArchitecturePreservingProvableRepair2023
0.7 Compiling Higher-Order Specifications to SMT Solvers: How to Deal with Rejection Constructively daggitt-2023-compiling
0.8 Vehicle: Interfacing Neural Network Verifiers with Interactive Theorem Provers daggitt-2022-vehicle
0.9 PRIMA: General and Precise Neural Network Certification via Scalable Convex Hull Approximations mullerPRIMAGeneralPrecise2022
0.10 Provable repair of deep neural networks sotoudehProvableRepairDeep2021
0.11 Tracking translation invariance in CNNs myburghTrackingTranslationInvariance2021
0.12 Towards Verified Artificial Intelligence seshiaVerifiedArtificialIntelligence2020
0.13 Verification of Deep Convolutional Neural Networks Using ImageStars tranVerificationDeepConvolutional2020
0.14 Stride and Translation Invariance in CNNs moutonStrideTranslationInvariance2020
0.15 Why do deep convolutional networks generalize so poorly to small image transformations? azulayWhyDeepConvolutional
0.16 Formal Verification of CNN-based Perception Systems kouvarosFormalVerificationCNNbased2018
Definition. Constant Presheaf constant-presheaf
The constant presheaf of a set on a category is a functor such that
I often call this the discrete presheaf for , but I donβt know if thatβs standard.
Finite Cardinal Arithmetic in a Topos finite-cardinal-arithmetic-topos
In a topos with a natural numbers object , you can define finite cardinals as objects that arise as the pullback along a morphism of the generic finite cardinal. Write for the cardinal corresponding to .
The above characterization is a little obtuse and does warrant some more explanation. One way to make it more concrete is that and , and that finite cardinals warrant a nice induction principle. If is a property expressible in the internal language such that satisfies , and that whenever satisfies then satisfies ; then every finite cardinal satisfies . That is, forms a -closed subobject of and thus is all of .
My working mental model is in a presheaf topos, where the natural numbers object can be defined explicitly as .
Definition 0.17. Constant Presheaf constant-presheaf
The constant presheaf of a set on a category is a functor such that
I often call this the discrete presheaf for , but I donβt know if thatβs standard.
The finite cardinals with respect to can be characterized then as,
These obey nice algebraic properties.
Definition. A Grammar of Finite Cardinals finite-cardinal-grammar
Define a grammar for each finite cardinal .
Then define the grammar of all finite cardinals as
Definition. Finite Unambiguity finite-unambiguity
Define a grammar to be finitely unambiguous if .
There is likely a better name for this.
Finite Unambiguity is Not Equivalent to Unambiguity finite-unambiguity-not-unambiguity
For a while I believed finite unambiguity to be equivalent to the other definitions of unambiguity.
Definition 0.18. Finite Unambiguity finite-unambiguity
Define a grammar to be finitely unambiguous if .
There is likely a better name for this.
Definition 0.19. Unambiguity as Subterminality unambiguity-as-subterminality
A grammar is unambiguous if the unique map into the terminal object is a monomorphism. That is, is a subobject of .
Definition 0.20. Unambiguity as Unique Map into Codomain unambiguity-as-unique-map
A grammar is unambiguous if for all grammars and maps we have .
In a category with terminal objects, this is equivalent to unambiguity defined via subterminality.
Definition 0.20.1. Unambiguity as Subterminality unambiguity-as-subterminality
A grammar is unambiguous if the unique map into the terminal object is a monomorphism. That is, is a subobject of .
We may use an analogy from the category of sets, however I had missed the infinite case when translating this idea to grammars.
It is true that if a grammar is unambiguous, then it is finitely unambiguous. However, the converse does not hold unless the grammar has finitely many parse trees for each string. This finiteness condition is a semantic one. If this proof of unambiguity were to be internalized, then it could maybe be captured through the lens of some grammar of finite cardinals.
Let . is finitely unambiguous but not unambiguous. The isomorphism between and amounts to building a bijection between and , which is straightforward.
What Connectives Preserve Finiteness? finiteness-preserving-connectives
In order to bridge the gap between finite unambiguity and unambiguity, we can try to restrict to grammars that have finitely many parse trees. Weβd expect this to address the concern semantically, as that fixes the problem when the parses are interpreted in .
We can define when a grammar is finite, and then try to prove that finiteness is preserved on some sane operations on grammars. For instance, , , and each preserve finiteness. I havenβt proven this myself (which would be a useful exercise), but the following is substantiated in any topos with a natural numbers object [johnstone-2002].
0.21 Finite Cardinal Arithmetic in a Topos finite-cardinal-arithmetic-topos
In a topos with a natural numbers object , you can define finite cardinals as objects that arise as the pullback along a morphism of the generic finite cardinal. Write for the cardinal corresponding to .
The above characterization is a little obtuse and does warrant some more explanation. One way to make it more concrete is that and , and that finite cardinals warrant a nice induction principle. If is a property expressible in the internal language such that satisfies , and that whenever satisfies then satisfies ; then every finite cardinal satisfies . That is, forms a -closed subobject of and thus is all of .
My working mental model is in a presheaf topos, where the natural numbers object can be defined explicitly as .
Definition 0.21.1. Constant Presheaf constant-presheaf
The constant presheaf of a set on a category is a functor such that
I often call this the discrete presheaf for , but I donβt know if thatβs standard.
The finite cardinals with respect to can be characterized then as,
These obey nice algebraic properties.
The case of is more problematic. Semantically, for all we may bound the size of the set of parse trees
Precisely knowing this bound isnβt too important, but certainly it exists. We could capture this behavior by just adding an axiom that if and are each finite, then is finite. Although, Iβd rather not add an axiom.
We could directly try to prove that preserves finiteness. The statement would follow from showing that preserves monomorphisms. That is, if and , it suffices to show that is a monomorphism. This would imply finiteness, because you may then apply this for and as . However, you would also need to show that is finite, which isnβt immediately clear.
Internal Finiteness internally-finite-grammar
Define a grammar to be finite (or perhaps subfinite) if it is a subobject of grammar of finite cardinals.
Where is defined as follows.
Definition 0.22. A Grammar of Finite Cardinals finite-cardinal-grammar
Define a grammar for each finite cardinal .
Then define the grammar of all finite cardinals as
Star Continuity is a Semantic Property star-continuity-semantic-property
Star continuity in Dependent Lambek Calculus can often be a convenient proof technique, but itβs important to remember that this shouldnβt be the first line of defense.
Star continuity holds in the Agda model, but does it hold in the syntactic model? I believe that it does because of the presence of the indexed coproducts. So perhaps it isnβt so sinister after all. It is worth noting that much of the reasoning performed by inducting on the length of a Kleene star isnβt very elegant. If a proof necessitates star continuity, then it doesnβt seem to be aided greatly by the type system.
Definition. Unambiguity as Subterminality unambiguity-as-subterminality
A grammar is unambiguous if the unique map into the terminal object is a monomorphism. That is, is a subobject of .
Definition. Unambiguity as Unique Map into Codomain unambiguity-as-unique-map
A grammar is unambiguous if for all grammars and maps we have .
In a category with terminal objects, this is equivalent to unambiguity defined via subterminality.
Definition 0.23. Unambiguity as Subterminality unambiguity-as-subterminality
A grammar is unambiguous if the unique map into the terminal object is a monomorphism. That is, is a subobject of .
Definition. Unambiguity via the Diagonal Being an Isomorphism unambiguity-via-diagonal
Finite unambiguity does not serve as an adequate definition of unambiguity that is equivalent to unambiguity as subterminality and unambiguity as a unique map. However, the definition attempted via finite unambiguity can be refined to something that is equivalent to these.
A grammar is unambiguous if is an isomorphism. Equivalently, is unambiguous if and are equal.
Definition. Subobject subobject
In a category , a subobject of is an isomorphism class of monomorphisms into .
Definition. Terminal Object terminal-object
An object in a category is terminal if there is a unique morphism from any other object into it.
Because terminal objects are unique up to unique isomorphism (as are all universal objects) we often just write to refer to the terminal object. Likewise, refers to the unique morphism into .
Definition. Equalizer equalizer
Let and be objects in a category with two parallel morphisms . The equalizer of and , if it exists, is the universal object with the following property:
- There is a morphism
Constructing Equalizers in Type Theory equalizers-in-type-theory
In the presence of -types, one may construct all equalizers. Given types and with functions , the equalizer may be constructed as
Definition. Subobject Classifier subobject-classifier
Subsets of a set may classically be identified with a characteristic map . Intuitively, for every , gives a truth value to the statement β is in the subset β. In this manner, the domain of the characteristic map, , classifies the subsets of .
Generalizing over this principle, in a category an object is a subobject classifier if maps into it from some object likewise uniquely identify a subobject of .
We can understand to behave like an object of truth values that are not necessarily boolean valued. A morphism can be thought of like a predicate on . If were a set, this would precisely be the characteristic function on it. However, this idea can generalize beyond sets. For instance in the category of graphs, is a cleverly constructed graph such that any graph homomorphism into it picks out a unique subgraph of .
There are always two (suggestively named) disjoint βpointsβ of , thought of as morphisms out of the terminal object . I think if weβre being careful, may properly be the βsubobject classifierβ but I always use the term to refer to itself.
A subobject induces a unique characteristic morphism such that
Moreover, the appropriate square must be a pullback.
Definition. Kleene Star in Dependent Lambek Calculus kleene-star
For a grammar , the Kleene star is defined as a least-fixed point,
Definition. Star Continuity star-continuity
A Kleene Algebra is star continuous if for all
Definition. Star Continuity in Dependent Lambek Calculus star-continuity-in-dependent-lambek
For a grammar , the Kleene star is isomorphic to an indexed coproduct.
That is, we may view the parses of like a linear list comprising parses of concatenated together. Further, for each of these lists we may know the precise length.
When viewing Dependent Lambek Calculus as a model of Kleene algebra, this is precisely the statement that star continuity holds.
Definition. First Set first-set
The first set () of a grammar are all the characters that may appear at the beginning of a word in the language of .
Definition. First Sets in Dependent Lambek Calculus first-set-in-dependent-lambek
The first set of a grammar may be captured in Lambek via the following proposition:
Or perhaps with ones of the grammars
Definition. FollowLast Set followlast-set
The followlast set () of a grammar are all the characters that may follow a word in the language of in a string that is in the language of .
Definition. FollowLast Sets in Dependent Lambek Calculus followlast-set-in-dependent-lambek
The followlast set of a grammar may be captured in Lambek via the following proposition:
Or perhaps with ones of the grammars
Definition. LL(1) Condition ll1-condition
A context-free grammar satisfies the LL(1) condition if it satisfies the following three conditions:
- All of its productions have pairwise disjoint first sets
- If a concatenation of nonterminals appears in a production, then has a disjoint followlast set from the first set of
- At most one production is nullable
This is essentially the type system of [1], which characterizes the LL(1) condition for context-free expressions.
Intutively, an LL(1) grammar can be parsed unambiguously, and without backtracking, by a predictive parser that only needs one token of lookahead.
Definition. Nullability in Dependent Lambek Calculus nullability-in-dependent-lambek
The nullability () of a grammar may be captured in Lambek via the following proposition:
Or perhaps with one of the grammars
Definition. Nullable Grammar nullable-grammar
A grammar is nullable if the empty string belongs to the language of .
Definition. Sequential Unambiguity sequential-unambiguity
Grammars and are sequentially unambiguous if the followlast set of is disjoint from the first set of .
We can understand this intuitively by characterizing the behavior of a left-to-right parser of . First it searches for a parse of , then upon finding a character that is not in it may begin trying search for .
That is, there is a unique boundary between the -parse and the -parse.
Definition. Closed Monoidal Structure closed-monoidal-category
A monoidal category is left closed if for each , the functor has a right adjoint forming the left internal-hom out of .
That is, for all , there is a natural isomorphism
There is an obvious right-handed variant that is right adjoint to .
If is both right and left closed, the monoidal category is simply called closed (or perhaps biclosed).
Am I on to Something? am-i-onto-something
Here I am documenting the half-baked ideas Iβve had but have never taken anywhere. On these Iβve either been out of depth on the requisite knowledge, rediscovered something that isnβt as new as Iβd hoped, or simply donβt have the time. I hope one day I can return to these and flesh them out.
2 entries
0.24 GrΓΆbner Bases for Inferring Polynomial Loop Invariants groebner-loop-invariants
When the tools in I4: Incremental inference of inductive invariants for verification of distributed protocols and On Symmetry and Quantification: A New Approach to Verify Distributed Protocols search for an inductive invariant of a distributed system, the search procedure instantiates a series of small finite models and tries to infer from their truth tables a series of logical formulae that hold over those finite models. These formulae are found by running the Quine-McCluskey algorithm for minimization of Boolean functions. The prime implicants found by Quine-McCluskey have a latent symmetry that can be abstracted into quantified formulae. There are only so many small numbers, and so small finite models may propose formulae that do not hold at larger sizes. However, if you find a formula that holds at size as well as size , then it is likely a good candidate to hold at all sizes. You need to be a little careful if your protocol is indexed by several variables, but mostly this general idea holds when abstracting to a protocol of unbounded size. Further discussion of this idea can be found in SAT-based quantified symmetric minimization of the reachable states of distributed protocols: An update.
My observation was that Quine-McCluskey is just a special instance of Buchbergerβs algorithm for computing GrΓΆbner Bases. That is, you can describe Boolean formulae as polynomials over the field with two elements, and in this translation Quine-McCluskey and Buchberger each compute the same data. This observation isnβt new in and of itself, but it does open up an opportunity to generalize the invariant search procedure that is used above.
The place I went looking to apply this idea was in the search of polynomial loop invariants. If a loop had an invariant that is expressible as a polynomial relation between the program variables, then you could apply the same idea as above to infer the loop invariant.
Suppose are the variables in scope of program and contains a loop that we want to infer an invariant for. Denote the value of at the -th loop iteration by . The invariant search procedure proceeds intuitively as the following: we will keep track of the minimal set of polynomials that could interpolate between all of the variable assignments that we have witnessed thus far. Formally this is kept track of by the ideal of polynomials. At the -th loop iteration we add a new generator to the ideal which corresponds to the assignments . The GrΓΆbner basis for this ideal provides the minimal data needed to generate all the assignments witnessed thus far, so if this process saturates then the GrΓΆbner basis encodes a polynomial loop invariant. The nice thing about polynomials is that they have finite degree which guarantees that this process does indeed saturate (provided that the degree of the invariant is smaller than the number of loop iterations).
I was so excited to find this idea. Iβd felt like it was my first good idea in grad school. Then I read Automatic Generation of Polynomial Loop Invariants: Algebraic Foundations and found out someone had done this 20 years ago. I still wonder from time to time if there is room to further refine this idea or perhaps further generalize it. For instance, there is further generalization beyond Quine-McCluskey or Buchberger to the Knuth-Bendix algorithm, which seems to be a more general instance of both of these algorithms. So perhaps this search procedure can be weakened to an even more general class? Although, Iβm not yet familiar much with the Knuth-Bendix algorithm.
0.25 A Method for Verifying Translational Invariance of Image Processing Neural Networks translation-invariance-verification
My understanding for how one may prove a safety property for a neural network is as follows. First, express as an input-output property. That is, choose some region in the domain of to represent the inputs of interest. Further choose some region in the codomain of to describe safe outputs. Then describe as the property:
For instance, this is more or less how Reluplex, Marabou, and -CROWN each work. Squinting my eyes, this is the only such method for verifying a safety property for . That is, this is the only way to get a 100% guarantee that holds rather than some high measure of confidence.
I believe in this formalism I have an idea for how to express the property β is translationally invariantβ where is an image classifying net.
To this end, we need a continuous artifact that captures what it means to translate an image. We can express an image as a matrix of pixels (or perhaps several parallel matrices if we care about color channels, but stick to a single grayscale matrix for now). To shift over by a single pixel, we may left-multiply by the shift matrix , where is the matrix filled with zeros and has 1β²s on the subdiagonal. Note that all the directions of shifting are captured by the combinations of left/right multiplication by , .
Shifting by multiple pixels is now expressed by the matrix , however this is still a discrete dynamical system. We donβt have a continuous object by which we can test our safety property. My initial thought to continuousify this system was to express some sort of exponential flow by . That is, consider the matrix
then captures what it means to translate the image over by , a real-valued βamount of shiftingβ. This choice of continuous artifact could then be used for a verification effort, however there are some issues related to numerical stability because isnβt invertible. Because it is not invertible, doesnβt actually exist. So to make the above construction work, we need to mildly perturb and take a pseudoinverse. This sort of works and makes it so does capture some real valued shift, but it is only accurate when is small. So this is maybe problematic for our verification effort. This may be resolvable by chopping up the problem into subproblems, each of which is thin enough for the current iteration is accurate enough on that subproblem. However, there are two big things to consider.
- The exponential approach above is probably too complicated. It seems
likely that you may be able to take some (sequence of) linear interpolation(s) between the βs. This will still be some continuous object that captures a real-valued shift and it will be much more stable than the approach above.
- Even if we sort out which continuous object represents translation
of an image, I cannot for the life of me train any neural network that is translationally invariant. So the proof method is useless if there is nothing that it would ever apply to.
Precisely in this last point, I have mostly focused on trying to train a small CNN for MNIST handwritten digit classification that preserves the output class for small, reasonable translations of the digit. Iβve used data augmentation to predispose the network to being translationally invariant, and even though I can get a high degree of invariance, I cannot get a network that is invariant for all of the examples even in the training set or a reserved testing set.
It may be the case that the method I propose for measuring this invariance could be used to adversarially train a network to have better invariance. It may also be the case that all CNNs are bound to suffer from small degrees of translational sensitivity. I cannot find the citation at the moment, but there was a paper that suggested that CNNs suffer from weird issues of translational sensitivity that relate to the size of the convolutional window. So maybe this approach is doomed to fail anyway.
On the whole, I will say that machine learning verification almost sounds like an oxymoron. That is, if you have the expressivity to properly state a sophisticated safety property, then you likely understand the problem enough to not need to resort to machine learning in the first place. So almost tautologically, it seems that there cannot be satisfying verification of neural nets, as the tasks of machine learning and verification live on very different epistemic foundations.
The related works I could liberate from Zotero may be found below.
15 entries
0.25.1 Adequate Losses via Quantitative Linear Logic capucci-2026-adequate
0.25.2 Quantitative Linear Logic for Neuro-Symbolic Learning and Verification flinkow-2026-quantitative
0.25.3 Quantifiers for Differentiable Logics in Rocq (Extended Abstract) marulandagiraldo-2025-quantifiers
0.25.4 Fundamental Components of Deep Learning: A category-theoretic approach gavranovicFundamentalComponentsDeep
0.25.5 Architecture-Preserving Provable Repair of Deep Neural Networks taoArchitecturePreservingProvableRepair2023
0.25.6 Compiling Higher-Order Specifications to SMT Solvers: How to Deal with Rejection Constructively daggitt-2023-compiling
0.25.7 Vehicle: Interfacing Neural Network Verifiers with Interactive Theorem Provers daggitt-2022-vehicle
0.25.8 PRIMA: General and Precise Neural Network Certification via Scalable Convex Hull Approximations mullerPRIMAGeneralPrecise2022
0.25.9 Provable repair of deep neural networks sotoudehProvableRepairDeep2021
0.25.10 Tracking translation invariance in CNNs myburghTrackingTranslationInvariance2021
0.25.11 Towards Verified Artificial Intelligence seshiaVerifiedArtificialIntelligence2020
0.25.12 Verification of Deep Convolutional Neural Networks Using ImageStars tranVerificationDeepConvolutional2020
0.25.13 Stride and Translation Invariance in CNNs moutonStrideTranslationInvariance2020
0.25.14 Why do deep convolutional networks generalize so poorly to small image transformations? azulayWhyDeepConvolutional
0.25.15 Formal Verification of CNN-based Perception Systems kouvarosFormalVerificationCNNbased2018
Definition. The Bicategory of Categories bicategory-of-categories
The bicategory of categories has
- as 0-cells, categories (at a fixed pair of universe levels, for objects and for morphisms);
- as hom-category , the functor category , so 1-cells are functors and 2-cells are natural transformations;
- as identity 1-cell, the identity functor;
as composition, on functors. On natural transformations and the horizontal composite is given directly by its components
Every component of the left unitor, the right unitor and the associator is an identity morphism, and their inverses are identities too. So the only content of the triangle and pentagon is that composites of identities are identities.
is still a bicategory and not a strict 2-category: and agree on objects and on morphisms, but in the formalization they are not the same functor definitionally. The structure cells are there to name that agreement.
A monad in is an ordinary monad on a category, and a prestack is a pseudofunctor into .
Definition. Category of Elements category-of-elements
Let be a presheaf on a category . The category of elements of is the displayed category over whose displayed objects over are the elements , and whose displayed morphisms over from to are proofs that .
Since is a set, there is at most one displayed morphism over each between given elements: a morphism of elements is a morphism of that happens to carry back to . Its total category is the classical category of elements , and a universal element of is exactly a terminal object of .
Definition. Corecursive algebra corecursive-algebra
Fix an endofunctor . An algebra is corecursive when for every coalgebra there is exactly one hylomorphism from to . Equivalently, the functor of the hylomorphism profunctor is constantly a singleton. It is the dual of a recursive coalgebra: an -algebra in is corecursive exactly when it is recursive as an -coalgebra in .
Example. If is a terminal coalgebra, then is invertible and is a corecursive algebra: a solution of is the same thing as a solution of , that is, a coalgebra map into the terminal coalgebra, and there is exactly one, .
Theorem. Day Convolution is Closed day-closed-structure
Let be a symmetric monoidal closed category that is complete and cocomplete. Let be a small monoidal -enriched category and be -enriched presheaves on . Define
Then, for Day convolution ,
Proof. Proof that Day Convolution is Closed day-closed-structure-proof
For , the enriched hom in is given by the end:
β
Symmetrically, and , so the enriched presheaf category is biclosed [1].
Proof. Proof that Day Convolution is Closed day-closed-structure-proof
For , the enriched hom in is given by the end:
β
Definition. Day Convolution day-convolution
Let be a symmetric monoidal closed category that is complete and cocomplete. Let be a small monoidal -enriched category. The Day convolution of -enriched presheaves is the enriched presheaf
with unit the representable .
Day convolution makes the enriched presheaf category a monoidal category, symmetric when is [1]. It is moreover closed.
Under the enriched Yoneda embedding the convolution of representables is representable, , so Day convolution is the cocontinuous extension of the tensor of .
As a Kan Extension
Equivalently, is the left Kan extension of along .
In
When , we recover the ordinary Day convolution of presheaves , where the formula simplifies to:
Definition. The Quotient and its Right Adjoint in Day Convolution day-quotient
Let be a small monoidal category. Recall that for presheaves , the Day convolution provides a closed monoidal structure:
which forms an adjunction .
For covariant functors and presheaves , we can define the quotient and its right adjoint, each of which is a presheaf on :
These form the adjunction .
Definition. The Quotient and Residual Coincidence in Day Convolution day-quotient-coincidence
The Day quotient, , takes to be a covariant functor and to be a contravariant one. While the residual takes both to be contravariant.
This apparent variance mismatch disappears when we restrict to the groupoid core , where variance is trivialized since .
For any functor on the core, we can freely extend it to both a presheaf and a covariant copresheaf on by left Kan extension along the respective inclusions and :
By substituting these extensions into the definitions of the residual and the quotient, we obtain a general coincidence for any and :
This equivalence states that computing the residual against the presheaf extension of is perfectly isomorphic to taking the quotient by its covariant extension.
The Representable Case
Instantiating the above theorem at a representable functor on the core, : the co-Yoneda lemma says that the left Kan extensions compute to the representables on :
Applying the general coincidence theorem, the left and right adjoints coincide precisely on (opposite-variance) representables:
When is the discrete monoidal category of strings, this recovers the derivative of formal grammars.
In nominal sets, I suspect that this construction also describes name abstraction and the freshness quantifier, although I have not check all of the details of the proof.
Definition. Dependent Tensor of Grammars dependent-tensor
The tensor of grammars,
is simply typed, so cannot see what matched.
Just as we have the generalization from the pair type to the dependent pair , we can define a dependent tensor operation. It is a -like generalization over the ordinary tensor, in which the second grammar depends on a parse tree of the first.
Write for the type of parse trees of , each paired with the string it parses.
For a grammar and a family of grammars , the dependent tensor is:
We may also write it as .
When is a constant family, the dependent tensor reduces to the ordinary tensor .
Dependence runs left to right here, which fits left-to-right parsing. The other handedness, in which the first grammar depends on a parse of the second, is also definable.
Relation to Day Convolution
Just as the ordinary tensor is given by Day convolution, I think there is a similar dependent Day convolution for which this operation is an instance. My best guess is for any presheaf on a monoidal category and a functor from the category of elements of into presheaves on ,
Iβm not positive on this though, and I would hope to also extend this to enrichments that arenβt but I donβt see how that could make sense given that I have quantified over the element in the coend rather than using the tensor of the enriching category.
Displayed Categories as Dependent Types displayed-categories-as-dependent-types
Displayed category theory is the category-theoretic analogue of dependent type theory. A category plays the role of a context, and a displayed category over it the role of a dependent type in that context. The analogy extends to each construction:
| Dependent type theory | Displayed category theory |
| context | category |
| dependent type | displayed category over |
| dependent function | section of |
| context extension | total category and its projection |
| substitution | reindexing |
| -type | displayed total category |
| a type not depending on its context | weakening |
Definition. Fiber Category fiber-category
Given a displayed category over and an object , the fiber of over is the category whose objects are the displayed objects , and whose morphisms are the vertical morphisms , those lying over the identity.
Identities are the displayed identities. The displayed composite of two vertical morphisms lies over rather than , so composition in the fiber transports it along .
Definition. Formal Grammar formal-grammar
For an alphabet , let be the free monoid of strings over . A formal grammar is a family of types indexed by strings:
This definition views formal grammars directly as indexed families of types over strings (which equivalently form presheaves on the discrete category of strings). This view forms the central notion of the Dependent Lambek Calculus. For a given string , the type represents the type of all valid parse trees for according to the grammar . If the grammar cannot parse , then is the empty type.
Definition. The Derivative of a Grammar grammar-derivative
When a grammar quotient is taken with respect to a representable grammar (matching exactly the string ), it coincides with the residual , as established by the Quotient and Residual Coincidence.
This special case is known as the Brzozowski derivative of by the string , often written as . We have the following coincidence:
Definition. Later on Grammars grammar-later
Write when is a proper suffix of , that is, with . The later of a grammar is
A parse of over is a parse of over every proper suffix of . In particular is a singleton.
This is later on families for strings under the proper-suffix order. That order is well-founded because it strictly decreases length, so it is a thin direct category. There is at most one map , so the product over maps from the strict past has one factor per proper suffix. Guarded recursion is modelled by presheaves on , the topos of trees, and more generally by sheaves over a well-founded base [1]. Here the later acts on families, which is what grammars are, and there is no clock.
Later is the right adjoint of the proper-suffix derivative. Let be the grammar of non-empty strings. Then the derivative has the right adjoint
so . Splitting into its summands gives the form of that the calculus can define:
The component at says: if the string begins with , then the rest parses as .
The restriction to non-empty is what makes this a later. At the component is . Including it would give a projection , and LΓΆb would then prove every grammar. For the same reason the later is a product over suffixes. A sum such as is empty at , and at it is . The identity step would then give LΓΆb a proof of .
In the Agda this is β· in Grammar/Later/Base.agda, which is defined as the indexed conjunction of βl-string. The mirror image β·r, over proper prefixes, is defined in the same way. See also the bilateral later and the later along an arbitrary well-founded order.
Theorem. LΓΆb Induction for Grammars grammar-lob
For every grammar and every term , where is the later on grammars, there is a unique global parse with
Here restricts a global parse to every proper suffix.
The proof is recursion on the length of the string. At , the parses already built at the proper suffixes of form an element of , and turns it into a parse at . Uniqueness is LΓΆb for families over the proper-suffix order. In the Agda, lob in Grammar/Later/Base.agda is this recursion, done by well-founded induction on length.
To prove an entailment this way, apply LΓΆb to . The hypothesis is the induction hypothesis at every proper suffix. It becomes usable once a non-nullable grammar has been consumed. If , then
because a parse of over splits with non-empty, so . This is β·-app-NE in Grammar/Later/Properties.agda. Induction on a Kleene star is the standard use.
Theorem. Next on Grammars Is Presheaf Structure grammar-next-presheaf
For presheaves, restricts along the strict past, as in later on presheaves. A grammar is only a family over strings, so a map into the later on grammars is extra data. It sends a parse over to parses over every proper suffix of .
Let , so that . This is the comonad for the suffix order, and presheaves are comonadic over families. The counit law forces the first component of a coalgebra to be the identity. Hence:
- a presheaf on strings under the suffix order is the same as a grammar with a map that satisfies coassociativity. Restricting to and then to must agree with restricting to directly;
- for a proposition-valued grammar (a language), coassociativity is automatic, and exists exactly when the language is closed under suffixes. For example, has one, but does not, since is empty.
LΓΆb does not need on . The fixed-point equation only restricts a global parse , and a global parse can always be restricted. In the Agda, the grammar-level IsCoalgebra record, with only the field next, is in Grammar/Later/Coalgebra.agda on the guarded branch.
On presheaves the strict downset of is represented by , because every proper suffix of is a suffix of . So predecessors simplify later to . On grammars there is no such simplification, since the factors at different suffixes are unrelated.
Definition. Grammar Quotients grammar-quotients
For formal grammars and , the quotient of by is the grammar of what is left of once an has been read off the front:
A parse of the quotient consists of an -parses at prefix and a -parse of the full string .
This operation has a right adjoint:
Intuitively, parses if: for all splittings of into a prefix and suffix, if the prefix matches then the suffix must match .
Together, these form the adjunction:
This is an instance of the Day quotient over the discrete monoidal category of strings.
Definition. Grammar Residuals grammar-residual
For formal grammars and , the residual (sometimes called the lollipop or linear implication) is the right adjoint to the grammar tensor . It describes strings that, when prefixed by a string matching , will match .
Formally, it is defined as:
Intuitively, a parse of at a string is a function that takes any prefix string and a parse of in , and produces a parse of the full concatenated string in .
Definition. Tensor of Grammars grammar-tensor
For formal grammars and , their tensor represents the concatenation of the languages they describe. A parse for at a string consists of a splitting of into a prefix and suffix , along with a parse of in and a parse of in .
Formally, it is defined as:
This operation is exactly the Day convolution of and , where we view formal grammars as presheaves over the monoid of strings.
Definition. Hylomorphism hylomorphism
Definition. The hylomorphism profunctor hylomorphism-profunctor
Fix an endofunctor . Hylomorphisms form a profunctor from coalgebras to algebras,
where is the set of with .
The action is by composition. If is a coalgebra morphism () and is an algebra morphism (), then is again a hylomorphism:
In these terms, a coalgebra is recursive when is the terminal functor, and an algebra is corecursive when is. For the inverse of an initial algebra the profunctor is representable, , and dually for a terminal coalgebra.
When every value of is a singleton, every divide-and-conquer specification over has exactly one solution. This is what local contractivity guarantees.
Definition. Lax Functor lax-functor
A lax functor between bicategories consists of
- a map on 0-cells, ;
- for all , a functor , acting on 1-cells and 2-cells;
- a unit comparison, natural 2-cells ;
- a composition comparison, 2-cells natural in and ;
such that three coherence laws hold, one for each structure cell of :
- left unit: ;
- right unit: ;
- associativity: .
Here between 2-cells is vertical composition, and and are whiskerings.
Lax functors compose: and . The coherence laws of the composite follow from those of and and naturality of , without using the triangle or pentagon of any of the bicategories involved.
A lax functor whose comparison cells are invertible is a pseudofunctor.
Definition. Lax and Pseudonatural Transformations lax-natural-transformation
Let be lax functors. A lax natural transformation consists of
- for each 0-cell , a 1-cell ;
for each 1-cell , a 2-cell filling the naturality square,
- such that is natural in : for a 2-cell , ;
- and such that respects the comparison cells of and : one law relating to , and the unitors, and one relating to , , , and the associators.
A lax natural transformation is pseudonatural when every is invertible. As with pseudofunctors, this is a property, so the pseudonatural transformations are a full subcategory of the category of lax transformations and modifications.
Between prestacks the 1-cells are taken to be pseudonatural. The reason is biuniversality: a transformation whose components are all equivalences of categories is an equivalence of prestacks only if its naturality cells are invertible, and a biuniversal element should be exactly a representation of a prestack up to such an equivalence.
Definition. Locally Discrete Bicategory locally-discrete-bicategory
Every category is a bicategory with only identity 2-cells. Its 0-cells are the objects of , and the hom-category is the discrete category on the set : a 2-cell is a proof that . Composition and identities are those of ; the unitors and associator are the unit and associativity laws of . Since the homs of are sets, any two parallel 2-cells are equal, so the triangle and pentagon hold trivially.
A functor gives a pseudofunctor , whose comparison 2-cells are the functor laws of .
The locally discrete bicategory is how ordinary indexed categories enter bicategorical language: a prestack on is a pseudofunctor , and its Grothendieck construction is a displayed category over .
Definition. Monoidal Category monoidal-category
A monoidal category is a category together with
- a functor , the tensor product;
- an object of , the unit;
natural isomorphisms
the associator, left unitor and right unitor;
such that the triangle and the pentagon below commute for all objects .
Definition. Nominal Sets nominal-set
Fix a countably infinite set of names . A nominal set [1] is a set equipped with an action by the group of finite permutations , such that every element has a finite support.
A finite set of names supports if any permutation fixing pointwise also fixes . The intersection of all supports for is called the least support, denoted .
The category of nominal sets is equivalent to the Schanuel topos. Under this equivalence, a nominal set corresponds to a functor , where is the category of finite sets and injections, given by mapping a finite set of names to the set of elements supported by :
Definition. Nominal Sets as Day Quotients nominal-sets-quotient
In the Schanuel topos, the underlying category for Day convolution is , where is the category of finite sets and injections.
Given a nominal set , its presheaf action describes elements supported by :
Instead of taking a priori as a presheaf, we can view it as a finitely- supported -set. We can restrict our attention to the groupoid core , asking for the support to be exactly the input:
This family is functorial on finite sets and bijections.
By extending this functor along the inclusions described in Quotient Coincidence, we can extend to both a presheaf and a copresheaf on . This suggests the equivalence:
I suspect this allows us to describe name abstraction [1] βwhich ordinarily looks like an operation on two presheaves of the same varianceβas the quotient of by the induced copresheaf .
Concretely, is usually given by the quotient of the product by an equivalence relation:
where for permutations fixing .
On the other hand, the Day quotient computes to a coend:
A priori, a coend over of the product is expressed as the quotient of a set of triples by an equivalence relation :
However, because the comprehension formula fixes , we reduce to a quotient of pairs:
I suspect that this quotient will equate to the one given by , thus resolving the apparent issues with variance. I further suspect that one will need the sheaf condition (pullback-preservation) of nominal sets to establish this equivalence.
Definition. Opposite Bicategory opposite-bicategory
The opposite of a bicategory has the same 0-cells and reverses the 1-cells but not the 2-cells:
Composition swaps its arguments, . The left unitor of is the right unitor of and vice versa, and the associator of is the inverse of the associator of , with its arguments reversed.
Since the 2-cells keep their direction, a lax functor induces a lax (not oplax) functor with the same action on cells. Reversing the 2-cells instead gives the bicategory , whose hom-categories are the opposites .
Duality saves work: a coherence lemma about can often be obtained by instantiating a companion lemma at , which swaps left and right.
Definition. Pseudofunctor pseudofunctor
A pseudofunctor is a lax functor whose unit and composition comparisons
are invertible 2-cells. So preserves identities and composition up to coherent isomorphism.
Being pseudo is a property of a lax functor: invertibility of a 2-cell is a proposition, since inverses are unique. The data of a pseudofunctor is exactly the data of a lax functor, and everything proved about lax functors applies to pseudofunctors unchanged. Pseudofunctors are closed under composition and identities, as lax functors are, because invertible 2-cells are closed under composition and under the action of a functor on hom-categories.
The main examples here are prestacks, pseudofunctors . When is locally discrete on a category , these are the pseudofunctors of fibred category theory.
Definition. Recursive coalgebra recursive-coalgebra
Fix an endofunctor . A coalgebra is recursive when for every algebra there is exactly one hylomorphism from to , that is, exactly one solution of
Equivalently, the functor of the hylomorphism profunctor is constantly a singleton.
Recursiveness is a coalgebraic form of well-foundedness: decomposes each input into subproblems, and recursiveness says that every divide-and-conquer program built on this decomposition has a unique meaning, without mentioning an order on inputs. [1] use recursive coalgebras on categories of indexed families to obtain algorithms that are correct by the type of the map they compute.
Example. If is an initial algebra, then is invertible (Lambekβs lemma) and is a recursive coalgebra. Precomposing with the isomorphism , the equation is equivalent to , which says is an algebra map out of the initial algebra; there is exactly one, .
The dual notion is a corecursive algebra.
Definition. The Schanuel Topos schanuel-topos
Let be the category of finite sets and injections. The Schanuel topos is the category of pullback-preserving functors .
Equivalently, it is the category of nominal sets, which are sets equipped with an action by the group of permutations on a countable set of names , such that every element has finite support.
Definition. Section of a Displayed Category section-of-displayed-category
A section of a displayed category over chooses
- for each object , a displayed object , and
- for each morphism , a displayed morphism lying over ,
such that and .
A section is to a displayed category what a dependent function is to a dependent type: it picks a displayed datum over every base datum.
Definition. Types over an Algebraic Theory theory-type
Fix a finitary algebraic theory : sorts , a signature , equations, and a set of generators. Write for the free -model on and for its carrier at sort .
A -type of sort is a family .
For -types and of sort we have the additives, defined pointwise,
along with , , and their indexed versions.
Each operation of gives a multiplicative, defined by Day convolution,
Each lifted operation also has closed structure: a residual in each of its arguments. Day algebras [1] give a more general convolutional definition.
When is the theory of monoids and is an alphabet, -types are formal grammars. Multiplication gives , the unit gives , and we recover Lambek.
Definition. Total Bicategory total-bicategory
The total bicategory of a displayed bicategory over packages the base and the displayed data together, one dimension up from the total category of a displayed category.
- Its 0-cells are pairs of a 0-cell of and a displayed 0-cell over it.
- Its hom-category from to is the total category of the displayed hom-category . So a 1-cell is a pair and a 2-cell is a pair .
- Identities, composition, unitors and associator are pairs of the base structure and the displayed structure over it, and the triangle and pentagon hold because they hold in the base and, over that, in the displayed bicategory.
Projecting to first components is a pseudofunctor whose unit and composition comparisons are identity 2-cells.