Reference. flap: A Deterministic Parser with Fused Lexing
Lexers and parsers are typically defined separately and connected by a token stream. This separate definition is important for modularity and reduces the potential for parsing ambiguity. However, materializing tokens as data structures and case-switching on tokens comes with a cost. We show how to fuse separately-defined lexers and parsers, drastically improving performance without compromising modularity or increasing ambiguity. We propose a deterministic variant of Greibach Normal Form that ensures deterministic parsing with a single token of lookahead and makes fusion strikingly simple, and prove that normalizing context free expressions into the deterministic normal form is semantics-preserving. Our staged parser combinator library, flap, provides a standard interface, but generates specialized token-free code that runs two to six times faster than ocamlyacc on a range of benchmarks.
Cite
Cites 34 works (5 here)
With notes (5)
A typed, algebraic approach to parsing krishnaswami_typed_2019
In this paper, we recall the definition of the context-free expressions (or µ-regular expressions), an algebraic presentation of the context-free languages. Then, we define a core type system for the context-free expressions which gives a compositional criterion for identifying those context-free expressions which can be parsed unambiguously by predictive algorithms in the style of recursive descent or LL(1). Next, we show how these typed grammar expressions can be used to derive a parser combinator library which both guarantees linear-time parsing with no backtracking and single-token lookahead, and which respects the natural denotational semantics of context-free expressions. Finally, we show how to exploit the type information to write a staged version of this library, which produces dramatic increases in performance, even outperforming code generated by the standard parser generator tool ocamlyacc.
Adaptive LL(*) parsing: the power of dynamic analysis parr-2014-adaptive
Regular-expression derivatives re-examined owensRegularexpressionDerivativesReexamined2009
Abstract Regular-expression derivatives are an old, but elegant, technique for compiling regular expressions to deterministic finite-state machines. It easily supports extending the regular-expression operators with boolean operations, such as intersection and complement. Unfortunately, this technique has been lost in the sands of time and few computer scientists are aware of it. In this paper, we reexamine regular-expression derivatives and report on our experiences in the context of two different functional-language implementations. The basic implementation is simple and we show how to extend it to handle large character sets (e.g., Unicode). We also show that the derivatives approach leads to smaller state machines than the traditional algorithm given by McNaughton and Yamada.
Deterministic regular languages bruggemannkleinwood
The ISO standard for Standard Generalized Markup Language (SGML) provides a syntactic meta-language for the definition of textual markup systems. In the standard the right hand sides of productions are called content models and they are based on regular expressions. The allowable regular expressions are those that are “unambiguous” as defined by the standard. Unfortunately, the standard’s use of the term “unambiguous” does not correspond to the two well known notions, since not all regular languages are denoted by “unambiguous” expressions. Furthermore, the standard’s definition of “unambiguous” is somewhat vague. Therefore, we provide a precise definition of “unambiguous expressions” and rename them deterministic regular expressions to avoid any confusion. A regular expression E is deterministic if the canonical -free finite automaton recognizing L(E) is deterministic. A regular language is deterministic if there is a deterministic expression that denotes it. We give a Kleene-like theorem for deterministic regular languages and we characterize them in terms of the structural properties of the minimal deterministic automata recognizing them. The latter result enables us to decide if a given regular expression denotes a deterministic regular language and, if so, to construct an equivalent deterministic expression.
Derivatives of Regular Expressions brzozowskiDerivativesRegularExpressions1964
Kleene’s regular expressions, which can be used for describing sequential circuits, were defined using three operators (union, concatenation and iterate) on sets of sequences. Word descriptions of problems can be more easily put in the regular expression language if the language is enriched by the inclusion of other logical operations. However, in the problem of converting the regular expression description to a state diagram, the existing methods either cannot handle expressions with additional operators, or are made quite complicated by the presence of such operators.In this paper the notion of a derivative of a regular expression is introduced and the properties of derivatives are discussed. This leads, in a very natural way, to the construction of a state diagram from a regular expression containing any number of logical operators.
External (29)
- flap: A Deterministic Parser with Fused Lexing (artifact) (2023)
- flap: A Deterministic Parser with Fused Lexing (arXiv) (2023)
- A practical mode system for recursive definitions (2021)
- ParTS: Final Report (2020)
- Proceedings of the 40th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI 2019 (2019)
- Sound, fine-grained traversal fusion for heterogeneous trees (2019)
- Generating mutually recursive definitions (2019)
- Push versus pull-based loop fusion in query engines (2018)
- Stream fusion, to completeness (2017)
- Optimizing CUDA code by kernel fusion: application on BLAS (2015)
- Staged parser combinators for efficient data processing (2014)
- Core_bench: Micro-Benchmarking for OCaml (2013)
- Towards Effective Two-Level Supercompilation (2010)
- Faster Scannerless GLR Parsing (2009)
- Stream fusion: from lists to streams to nothing at all (2007)
- Context-aware scanning for parsing extensible languages (2007)
- Common Format and MIME Type for Comma-Separated Values (CSV) Files (2005)
- Packrat parsing: simple, powerful, lazy, linear time, functional pearl (2002)
- Disambiguation Filters for Scannerless Generalized LR Parsers (2002)
- Greibach Normal Form Transformation Revisited (1999)
- Multi-Stage Programming: Its Theory and Applications (1999)
- Eta-expansion does The Trick (1996)
- Deforestation: transforming programs to eliminate trees (1990)
- Compilers: Principles, Techniques, and Tools (1986)
- How to Replace Failure by a List of Successes (1985)
- Strict Deterministic Grammars and Greibach Normal Form (1979)
- Normal forms of deterministic grammars (1976)
- A New Normal-Form Theorem for Context-Free Phrase Structure Grammars (1965)
- The Menhir parser generator (software)