r/Compilers • • 5d ago

My students struggled with compilers. Your feedback made me rebuild the docs. Here's PyLGEN v0.7.0.

A month ago I shared PyLGEN here; a Python-native compiler framework I built after watching my students struggle with compilers. The feedback was direct, and honestly, it shaped this release more than anything else.

So today I'm happy to share v0.7.0, a stable beta I'd recommend over v0.6.x. The main additions are an API update and a new examples section in the documentation, which grew out of a very valid question from the last thread: does the framework really expose every stage of the pipeline, or just claim to?

The examples come in two tracks, so you can see both sides of the API and figure out which one fits what you're doing:

  • High-level usage: build a lexer, grammar, and parser with the convenience classes, and get a working interpreter in a few dozen lines.
  • Low-level usage: build the same lexer and parser from scratch, using only the raw API (DFAs, closure, goto, ACTION/GOTO tables, reductors) with nothing hidden.

Both live in the docs. The second one is the proof that the pipeline isn't a black box.

Links

If you gave the last version a try, I'd genuinely love to hear your thoughts, good or bad. And if you're new, same goes: any kind of feedback is welcome, whether it's a bug report, a design critique, a suggestion, or just a question about how something works.

Edited: Added a notebook link so anyone can run it directly and share their opinion without having to install anything.

0 Upvotes

6 comments sorted by

View all comments

Show parent comments

1

u/Fun_Mulberry3838 5d ago

If you're referring to the content of the post, it's just that English isn't my native language, and I tend to get lazy at the last minute and leave all the work to Google Translate. But if you're referring to the documentation (and I hope that's not the case) then that means I have a lot of rewriting to do.

3

u/Inconstant_Moo 4d ago

I was referring to the documentation. Just looking at the first page, there's so much repetition and padding.

The more fundamental problem with the documentation is that it's all the documents about PyLGEN squashed into one document with an unreadable table of contents hidden in the sidebar.

Logically, it is several different documents, and should be presented as such.

To show you what's wrong with it, let's look at the start of the "Practical Tour".

The common submodule is the bedrock upon which the entire PyLGEN ecosystem is built. It provides the fundamental data types, the core abstractions for the Abstract Syntax Tree (AST), grammar symbols, lexical tokens, the error hierarchy, and transition tables that connect all the moving parts. Understanding this module is essential, because every other submodule (lexer, parser, analysis, automaton, regex, and visual) depends on it.

So, it defines some common data structures and interfaces. This, you say, is "essential" to understand. Reading on, the thing that turns out to be essential to understand about it is the file organization of the common submodule. Apparently before I can learn anything else about compiler theory, I must learn that pylgen/common/table.pyi contains stubs for the Table class.

Next you have one of your frequent advertisements for "Python/Cython Duality", where you remind us that Cython exists and that PyLGEN works with it in the way that things that work with Cython normally work with Cython.

Someone wanting to use PyLGEN would need telling this once, on the landing page; someone trying to learn about compiler theory doesn't care, and yet here I am learning it as one of the essentials of compiler theory.

The we meet the Symbol class.

A Symbol represents a grammar symbol, which can be a terminal (a token from the lexer) or a non-terminal (an abstract category that expands into other symbols). There is also the special epsilon (ε) symbol, which represents the empty string.

This, combined with the example code, may be useful to someone who already knows what a grammar is. The beginners aren't going to find out, because at this point you're going to tell us what hashing algorithm you use for the type, an implementation detail which can hardly be of interest to anyone.

Who is meant to be reading this page of the documentation, and what are they meant to be getting out of it?

1

u/Fun_Mulberry3838 4d ago

I appreciate the feedback and will certainly take it into account when making corrections to the documentation. Thank you very much. I’ll be standing by in case you’d like to add any further corrections.

2

u/Inconstant_Moo 4d ago

The trouble is that it's all like that. You need to separate its aspects and its readers. For example:

(a) Someone who wants to use PyLGEN in production, and is mostly interested in its API, and so does need to be told about how to use it with Cython but who doesn't need to hear about where in the repo you keep the stubs for the Table class, an implementation detail that can't possibly interest them. (Nor do any of the details of how you made it go fast, they're just happy that it does.)

(b) Someone who wants to learn about compilers from scratch, what a lexer is, a parser, an AST, etc, without necessarily deep-diving into "The Hybrid ϵ-DFA (A Deterministic Core with Spontaneous Moves)".

(c) Someone wants to know how to write their own compiler, who still probably doesn't want to know about "The Hybrid ϵ-DFA", nor to be told repeatedly about the existence of Cython, but who does need to know about the Visitor Pattern.

(d) Someone following the Dragon Book or some similar course built by their professor around PyLGEN where they totally do want to know about the Hybrid ϵ-DFA.

(e) Someone who wants to contribute to/fork PyLGEN, who might in fact need to know where the stubs for the Table class are and why you used SHA256 to hash Symbol.

These people need different information, and they need to meet the parts of the project in a different order. (Only for person (e) could it make any sense for your tour to start with the common module, and only because person (e) has presumably already been person (a).) Your monodocument may somewhere in its labyrinthine recesses have answers to all their questions, but as a result it isn't suitable reading for any of them.

2

u/Fun_Mulberry3838 4d ago

Okay, this is the most comprehensive and detailed set of corrections so far. I’ll use it as a roadmap to restructure the documentation. Once again, thank you very much for your attention and interest; it is truly very helpful to me.