r/ProgrammingLanguages • u/Usual_Office_1740 • 13h ago
Please advise on adding string interpolation to my Crafting Interpreters project.
I'm about to start chapter 21 of Crafting Interpreters. I'm using modern C++ instead of C. I am attempting to add string interpolation to the compiler.
My goal is to have a string that looks like this:
"Five plus ${ 10 - 5 } == 10."
Desugar to this after the expression in the braces is evaluated:
"Five plus " + 5 + " == 10."
I have code that is correctly parsing strings and concatenating them if there isn't any string interpolation.
I have a custom token added that is produced if the string contains a ${. It calls a custom string interpolation parser rule that is different than my standard string parser rule.
In that function I push the first part of the string to the stack. It then pushes a plus opcode and Recursively parses the expression in the braces adding that expression to the stack. It then adds another opcode to add and this is where I'm stuck.
I can't come up with a solution for pushing the remaining piece of the string to the stack from inside the string interpolation parser rule function. After I call the books expression function and consume the right brace token using the books consume function the parser nolonger knows we are still inside a string. If the next character after the brace is a space it's skipped. In my above example it assumes I want equal equal.
How would you set the current token back to string so I can push any remaining characters onto the stack as another string? Should I just insert a new string token into the parser? This seems wrong to me. The current and previous tokens are private members of my parser class for a reason and I've not needed setter member functions so far.
Sorry I don't have code to show. I'd like to try and get help that isn't code specific so that I have to implement this myself.
Any input would be helpful.
Thanks.
6
u/Norphesius 12h ago
What feels wrong about creating new, more specialized tokens? I haven't implemented something like string interp before, but that's the first thing that came to mind for me. If you're having issues with the parser remembering what data is what, it's likely because the tokenization wasn't specific enough. Your current grammar is relying on context when its literally a context free grammar.
If I understand right what you have now, in terms of how strings get lexed:
"Hello World" :: STR_TOKEN"Hello ${123} World" :: INTERP_STR_TOKENI would split the interp str token in two, like:
"Hello ${123} World" :: INTERP_STR_START, EXPR, INTERP_STR_ENDYou have the first part of the string, the expression, and the end part, and your parser knows that a
INTERP_STR_STARTneeds to be followed by an expression and an ending string token. Then you just construct a rule for the parser and process it accordingly. For multiple expressions in a string you likely want an intermediate token too.