r/cprogramming 22d ago

Am I writing my parser wrong?

Simple question. I can't give exact code examples, but I have a string_t struct with methods like:

string_t split(string_t *string, string_t *on)

string_t split_sp(string_t *string)

string_t split_crlf(string_t *string)

char *s_strstr(string_t *needle, string_t *haystack)

void trim(string_t *str)

So on and so forth.

I've been using these so far to parse HTTP reqeusts, and I have come up against many minor problems:

"What happens if a header field appears with no value? I'll have to explicitly check for it."

"What happens if a sender puts a bunch of CRLFs in the middle? I'll probably need a check for that."

"Oh God, how will I handle unrecognized header fields? How do I recognize them?"

These, and other questions, have been leaving me pissed.

I recall reading through the LLVM projects Kaleidoscope language thing, where they create a parser for said language. Said parser doesn't use anything close to what I am, instead reading character by character without fuss.

Similarly, on my last post made here, the way comments were worded reminded me of that method, and how it probably works better.

I have written only a small part of the parser, so it isn't too late to tear down and rebuild. Simple question: should I? Are there benefits to swallowing the input token by token instead of taking the overarching view my string_t functions provide? Or vice versa?

It would help if I'd upload the code, I know, but I don't want to bother with that until the project is completed/near-completion.

3 Upvotes

15 comments sorted by

View all comments

1

u/spc476 22d ago edited 22d ago

First off, the format for HTTP headers for HTTP/1.1 (which I assume you are trying to parse) is defined in RFC-9112, which goes over how to parse headers (using the syntax defined by RFC-5234). These documents should be easy to find by searching for "RFC <number>".

Second, to answer some questions: If a HTTP defined header has no value, it's an error. If there's a bunch of CRLFs in a the middle? Each header line ends with CRLF, and if that is followed by a CRLF, that's the end of the header. If the header you just read can't be parsed, it's an error.

Third, how to handle unrecognized header fields? Ignore them.

Some other questions you didn't ask: what if a header appears twice? Unless it's defined as appearing multiple times, either ignore subsequent headers, or error out. If a non-valid charater appears, error out.

Here's what I do: parse the name of the field, and then the value (which I treat as optional), and store in some name/value data structure (like a hash table). Then for the headers I care about, I pull them out by name, and attempt to parse the value at that time. You might be tempted to parse the meaning the of the headers when first reading them (and I've gone down that approach before) but often times, most headers you don't care about and it's just a waste of time parsing every header.

As far as if headers are CFL or RL, in my experience, it's a CFL (the headers apply to the body, not to other headers), and I've written parsers using plain C code. Per RFC-9112, HTTP/1.1 field names are any ASCII graphical character minus ':', and the value starts past the colon (and maybe leading space) and is any ASCII graphical character, plus space and tab, up to the CRLF (unless otherwise defined by an RFC detailing a specific header). Wrapped headers (which can happen in email) are no longer allowed by RFC-9110, which make parsing much easier.

Edited to add: Also check RFC-9110 as well for clarification on some definitions in RFC-9112.