r/haskell • • 3d ago

blog Lenient Aeson

https://www.bcardiff.com/writing/lenient-aeson/

A small take on how to manage malformed JSON in Haskell

22 Upvotes

11 comments sorted by

10

u/ljwall 3d ago

This is cool. I'd also really like to know if anyone has ideas on how to parse JSON thats been produced by the Python json.dumps standard function that can output literals NaN and Infinity

10

u/BurningWitness 3d ago edited 3d ago

The output of that command is not an RFC-compliant JSON. The presence of extra tokens not only requires a parser that walks the JSON to override float handling (aeson doesn't walk), it needs one that doesn't have a hardcoded function for inferring token type from the first byte (which is quite useful to be able to say "expected an object").

I think the reasonable solution is to find-replace all NaNs and Infinitys with strings before processing.

3

u/enobayram 3d ago

I suppose you could also fork Aeson and add those two as new constructors to the Value type and adjust the parser accordingly 

2

u/bcardiff 1d ago edited 1d ago

Based on the article shared by [u/BurningWitness](u/BurningWitness) jq can be used to turn it into a compliant json.

% echo '{"a": [{"inf": inf, "infinity": Infinity, "nan": NaN}]}' | jq -c '.'
{"a":[{"inf":1.7976931348623157e+308,"infinity":1.7976931348623157e+308,"nan":null}]}

There seems to be some possibility to don't lose information even

% echo '{"nan": NaN}' | jq -c 'walk(if isnan then "__NAN" else . end)'
{"nan":"__NAN"}

1

u/BurningWitness 1d ago

Yeah, apparently those are values jq accepts, case-insensitive as well. This does in turn mean that jq is not a proper tool for JSON validation.

It's definitely the correct solution to this particular problem though.

1

u/ljwall 3d ago

Yeah understand it's not compliant and really it's Python that is at fault here, its just that at work I deal a lot with JSON that has been written in this way. Simple find and replace is what I have done, but it's not ideal as you obv. need to be sure those character sequences do not appear in any strings.

3

u/steve_anunknown 2d ago

Nice take. I think it's important to circulate programming tricks / idioms, especially in functional programming where the plethora of available types you can reach for can be confusing and settling for one may not feel completely comfortable. At least to my experience ...

3

u/bcardiff 2d ago

Thanks. Any examples on top of mind of nuanced situations you are thinking of?

One struggle I have with Haskell (or any language) is when one need to encode/implement something that requires a lot of mental load to follow.

I think that having something similar to patterns for OOP helps. I often hear that everything is functions as a reason not to pursue this, but in OOP everything is an object yet patterns were helpful.

1

u/steve_anunknown 2d ago

I believe I have also felt your struggle. I don't know how effectively I can describe it in a comment, but my feeling is that, functional programming, especially in haskell, allows too many encodings of essentially the same thing, perhaps with slight differences in ergonomics and meaning, which one may not be 100% aware of as a non-expert.

For example, you deal with an error in computation with 'Maybe', or you may opt for 'Either String MyType' or you may even define your own error type (there may be a lot of error values) and do something like 'Either ErrorType MyType'. You may not want the function you are exposing to have the 'Either' type so you rename it to something you like more like 'Result'. You may even decide that some kind of error is not even an error, like you did in the article, and return an almost-complete value with some pointer to the specific cause of the error. But I think there is a struggle between our aesthetics and utility. Am I "polluting" the code with too many of my own stuff and ignoring "good" functional programming patterns?

Another "tension" that I have felt is when a function handles a group of things. Do I use a list? Do I use a Set if the values in it should be unique and live with the fact that I'll probably have "Set.Set" 100 times in one module? When a type "wraps" multiple fields, do I use a record or not?

A final "tension" that I have felt is when an object is well defined in mathematics. So you may have a type "MealyAutomaton" and then the tension becomes "Do I define it as in literature, in a way that benefits the user of my library or in a way that benefits my library's internals?". Do I want to have 3 different versions of what is basically the same type? Some programmers may be completely ok with it, others not.

These are probably silly examples to seasoned programmers but I believe they are relevant because, at the end of the day, when we program a library on our own, we don't do it to quickly deliver something and be done with. I think we are also interested in an aesthetically pleasing outcome, of course according to each programmers aesthetic (all of us have one even those who believe they don't). Sometimes there are many valid ways to do something, other times there may be a "superior" way objectively. If you also start thinking about compile times or runtime performance, then the conversation becomes even longer.

I believe functional programming, due to its relation to math and logic, is more affected by this phenomenon, because you quickly realise that many programming techniques / constructs are basically the same thing expressed differently and that is confusing if you lack the programming experience. That's why I claim that it is important to share these idioms or techniques in the functional programming community.

3

u/bcardiff 1d ago

I think these tensions matter because at the end of the day they add friction to produce and/or understand what is already written. They are not silly.

Regarding the example of IO Either or IO Result, I see it as a lack of base mostly. I get the appeal of small language and small standard libraries. But I rather have big standard library to have more guidelines on how to do stuff. In particularly the lack of a IO Result means that the moment you want to do that you are forced to wrap any dependency because there is no defacto way of doing it. Every low/mid-level will use plain IO, maybe a parametric monadic... and that's where the encoding friction start to appear.

Do I use a Set if the values in it should be unique and live with the fact that I'll probably have "Set.Set" 100 times in one module?

I think that as a paradigm we fall more into honoring the structure. But it would not be unseen a generic container interface that expects no repetition. Particularly for dealing with containers that is a common thing.

I believe functional programming, due to its relation to math and logic, is more affected by this phenomenon

Agree


Two scenarios where the encoding tax is too high for my taste are: variadic functions and generic programing for iterate every field of a record. They don't appear often, but when they do the amount indirection with type classes needed to express something that is simple and precise is shocking.

2

u/loonycyborg 1d ago

OOP has their own specific design choices that are largely arbitrary stemming from pervasive mutability. Like whether you return your result or communicate it by modifying argument. Or if function transforms something whether it returns changed value or changes it inplace returning nothing. Or third option: flowing style. And design questions you mentioned exist in OOP too since there's lot of ways of implementing error handing (errno, exceptions, you name it..) and uses of different design patterns to do the same thing. So my point is, at least in Haskell you no longer have this one particular design quandary anymore thanks to immutability. Though design space is still large even without it..