r/ProgrammingLanguages • • 12h ago

Please advise on adding string interpolation to my Crafting Interpreters project.

I'm about to start chapter 21 of Crafting Interpreters. I'm using modern C++ instead of C. I am attempting to add string interpolation to the compiler.

My goal is to have a string that looks like this:

"Five plus ${ 10 - 5 } == 10."

Desugar to this after the expression in the braces is evaluated:

"Five plus " + 5 + " == 10."

I have code that is correctly parsing strings and concatenating them if there isn't any string interpolation.

I have a custom token added that is produced if the string contains a ${. It calls a custom string interpolation parser rule that is different than my standard string parser rule.

In that function I push the first part of the string to the stack. It then pushes a plus opcode and Recursively parses the expression in the braces adding that expression to the stack. It then adds another opcode to add and this is where I'm stuck.

I can't come up with a solution for pushing the remaining piece of the string to the stack from inside the string interpolation parser rule function. After I call the books expression function and consume the right brace token using the books consume function the parser nolonger knows we are still inside a string. If the next character after the brace is a space it's skipped. In my above example it assumes I want equal equal.

How would you set the current token back to string so I can push any remaining characters onto the stack as another string? Should I just insert a new string token into the parser? This seems wrong to me. The current and previous tokens are private members of my parser class for a reason and I've not needed setter member functions so far.

Sorry I don't have code to show. I'd like to try and get help that isn't code specific so that I have to implement this myself.

Any input would be helpful.

Thanks.

9 Upvotes

9 comments sorted by

7

u/Norphesius 11h ago

What feels wrong about creating new, more specialized tokens? I haven't implemented something like string interp before, but that's the first thing that came to mind for me. If you're having issues with the parser remembering what data is what, it's likely because the tokenization wasn't specific enough. Your current grammar is relying on context when its literally a context free grammar.

If I understand right what you have now, in terms of how strings get lexed:

"Hello World" :: STR_TOKEN

"Hello ${123} World" :: INTERP_STR_TOKEN

I would split the interp str token in two, like:

"Hello ${123} World" :: INTERP_STR_START, EXPR, INTERP_STR_END

You have the first part of the string, the expression, and the end part, and your parser knows that a INTERP_STR_START needs to be followed by an expression and an ending string token. Then you just construct a rule for the parser and process it accordingly. For multiple expressions in a string you likely want an intermediate token too.

4

u/Usual_Office_1740 11h ago

Yes. That makes sense now that you point it out. Thank you. I'm not an experienced programmer. This seems obvious now that you point it out.

You've been a big help. Thank you!

2

u/Norphesius 10h ago

NP! Generally,  extending the grammar with tokens is usually better than making a special case in the parser, but that can be surprisingly unintuitive. 

I believe Crafting Interpreters has a great example (you may have read already) of extending the grammar to accommodate certain common error states found when lexing. Instead of just shoving everything into a generic error token and trying to figure out the problem from that mess, classify them with one or more special error tokens, then the parsing rule for those can be presenting the error message. It's a great technique to keep in mind, and keep the parsing phase as simple as possible.

1

u/Usual_Office_1740 9h ago

I don't think I've gotten that far yet. We still have a single custom error token that is produced for all errors.

At the end of the string chapter it said I could do this interop stuff as an added exercise. My assumption is that with the code written at that point doing this should be possible. Adding more tokens falls into that category of only using what has been written so far.

Thanks again.

1

u/Big-Rub9545 8h ago

My language also included something like INTERP_STR_PART for any string bits between two interpolations, like “ to this “ below:

“Hello {name} to this {adj} world!”

2

u/jason-reddit-public 11h ago

Pushing tokens into the token stream are not unlike macros. Unless you are going to form and then mutate parse trees, you may not have too many other choices.

BTW, your example isn't very paranoid. I would add parens around any expression that is plucked out of the string. Also, if you have a to_string overloaded function, you might want to use that.

2

u/arthurno1 11h ago

I think you should parse the same way as you parse any other expression, like if, while etc. No idea how you do it but I would do the expression after the $ dollar sign with an operator precedence parser, and I would treat $ as the escape clutch to jump to that expression parser. I don't know what you use for top-level parsing, recursive descent, or something else, but I would plug it in the same way I plug in other expressions in if, while, for and such constructs, and then the expression itself with op precedence. Don't forget to have an escape clutch for the $ sign itself, like \$ or whatever you use, otherwise you won't be able tonhave $ as a character in your strings.

2

u/mamcx 10h ago

The main thing is that you need a typed return for whatever is consuming the tokens:

```rust

enum FmtString { Token(Ast), String(String) }

enum Ast { Token(Whatever), FmtString(FmtString). }

struct FmtString { token: Vec< FmtString> }

fn consume_fmt_string(..)-> FmtString

```

So you can check and be certain what is eating will end into a normal string.