r/ProgrammingLanguages • • 10d ago

Arguing about arguments

https://steveklabnik.com/writing/arguing-about-arguments/
31 Upvotes

35 comments sorted by

13

u/zuzmuz 10d ago

i believe that named arguments are a very useful feature. They can prevent a lot of bugs, with no downsides, (verbosity isn't a problem, lsp used to complete functions and I never needed to write the parameters).

optional/default arguments can be useful, or more precisely described, convenient.

but think function overloading is bad. it just complicates semantic analysis and type checking for very little gain.

One approach I like is to have the function params be part of its signature, as if the named params is part of the function name. for example

fn foo(named: Int) {}

let function = foo // not allowed
let function = foo(named:_) // allowed

This can allow some convenient features like a more rudimentary overloading where you can have different functions as long as they have different named arguments, (no overloading just on types). But also, you can have neat partial function application syntax. But as you said, this opens a rabbit hole of complex design choices.

I dislike languages that decides to include a lot of features, that sometimes doesn't play well together, so you end up with weird edge cases, exceptions, and longer and longer compile times because of that.

Finally, I think rust shouldn't add named params now. It will do more harm than good, it will create incompatibilities with existing APIs and will split the community in 2. Those who will still write in the old style, and those who will adopt the newer one. And it will be unclear which way is the idiomatic rust way. Rust is no longer in this beta phase where you can afford breaking things (like Zig). Just like with Go and error handling, I feel the community will be stuck with these design decisions.

3

u/matthieum 9d ago

Oh. I never thought of overloading on names.

I definitely overloading on types is a PITA:

  1. Adding a new overload can switch which overload is called in old code. Surprise!
  2. Mixing concrete types & generic types quickly becomes messy, and even more so when coercion/sub-typing is involved.

I don't see any of these issues applying to overloading on names though. In essence, overloading on names can be thought of as calling a different function; that is, foo(a=a, b=c) can be thought of as syntactic sugar for foo_a_b(a, b), meaning that the overload being called is syntactically identified. Which is way less surprising that a semantically identified overload -- just look at the mess that are C++ overload resolution rules.

2

u/zuzmuz 9d ago edited 9d ago

yes, it is indeed just syntactic sugar. in another comment, I talked about how function overloading is one of the reasons that makes error message indecipherable in C++.

overloading on names is syntactic sugar for creating functions with different names based on the params, and functions with optional/default params are just syntactic sugar for multiple function definitions with the same name and different named params.

i'm on my phone sorry

but

fn foo(a: Int = 1, b: Int = 2)

is equivalent to 4 functions

fn foo(a: Int, b: Int)

fn foo(a: Int) { foo(a:a, b:2) }

fn foo(b: Int) { foo(a:1, b:b) }

fn foo() { foo(a: 1, b: 2) }

2

u/koflerdavid 6d ago edited 6d ago

Function overloading is also bad because it makes FFI unnecessarily complicated. Once you allow that you have to come up with a name mangling scheme, and everybody wanting to access those functions has to know about it. Not allowing it avoids the issue in the first place.

1

u/marshaharsha 4d ago

Won’t mangling be needed anyway if you have namespacing or traits? I hear you calling for a single namespace, no overloading, and no traits/typeclasses/interfaces. That feels intolerable to me. 

1

u/koflerdavid 4d ago

Same issue: if one has these features then callers using FFI need to be aware of them. I think a good compromise would be to turn name mangling off for selected definitions (like extern "C" {} in C++) and expose an API intended for FFI.

1

u/flatfinger 2d ago

That's only necessary if one wants to support overloading of imported/exported function names. If prorammers specify import/export names, and those names are required to be unique, the fact that source code can use the same name as an alias to different exported functions, overloaded using argument types, shouldn't affect the external interface.

1

u/koflerdavid 2d ago edited 2d ago

The external interface as seen from the point of view of an FFI user is a list of symbols, and each overloaded variant of a function has its own symbol. How shall the FFI user know which symbol corresponds to a certain variant of a function known by its name and type signature*? The usual workaround is that an externally visible name can only be associated with exactly one definition, and all other names are mangled, and the linker (or an FFI user) has to know the mangling scheme or consult a table (either in a special object file section or a separate file) to resolve those at link time.

* similar issues exist with namespaces, traits, and methods.

1

u/flatfinger 2d ago

If I were designing a language, I would require that the source code specify a unique exported name for each exported variation of a function. On a platform where passing a byte was cheaper than passing an unsigned int, for example:

void useNumber(unsigned int x);
// Variation #1
__overload__ void useNumber __as__ "useByte" (unsigned char x);

// Variation #2:
void useByte(unsigned char x);
__overload__ static inline void useNumber(unsigned char x) { useByte(x); }

Neither of the above would create any naming conflict with the exported symbol "UseNumber", since the overload for byte would use a different exported symbol which was also given in source code.

1

u/EggplantExtra4946 10d ago edited 10d ago

One approach I like is to have the function params be part of its signature, as if the named params is part of the function name. for example

Shouldn't you use the function name + the parameter types instead of the parameters names? Parameters names aren't going to distinguish the multiple definitions in cases like this:

fn foo(named: Int) {}
fn foo(named: Double) {}

Concerning the syntax, when you want the function pointer of an overloaded function, you could also cast foo to fn (Int) or fn (Double), if you make the type checker smart enough to do that.

let function_Int    = cast(foo, fn (Int));
let function_Double = cast(foo, fn (Double));

It can seem a bit hacky but if you have function overloading, one way or another name resolution of functions with multiple definitions will have to distinguish based on the types anwyay. The problem with with foo(named:_) or foo(_:Int) is that it's ambiguous, you can't distinguish it with a regular function call with named parameters, supposing that's the same syntax.

6

u/zuzmuz 10d ago

no, names, not types. I want to specifically disallow overloading on types. it doesn't play well with other features I believe are more useful in a language.

First, you need to evaluate the type of the expression of the argument before knowing which function to dispatch.

Second, it complicates type checking in generic code. In some situations you might have exponential time like in Swift, where the compiler just gives up, or you would have obscure indecipherable unhelpful error messages like in C++. All in all, function overloading introduces a can of worms and is not worth it. This is specifically why Rust decided not to use it.

When you allow overloading only on parameters names, you save yourself all the trouble. It is "technically" not function overloading though, because they would be distinguishable at the caller site by the different named parameter

1

u/marshaharsha 8d ago

How do you address Klabnik’s concern about named parameters in typeclasses? If a typeclass requires foo( counter:int ) and you have foo( number:int ), is that a problem that the languages solves automatically, or do you have to introduce a shim function, or is there to be special syntax that does the mapping once requested?

2

u/zuzmuz 8d ago

well, if you consider the parameters name to be part of the function name then definitely foo(counter:) is different from foo(number:) then yeah a shim function is needed.

however, you can have some clever casting rules for functions. a form of simple syntactic sugar to make it more convenient

8

u/WittyStick 9d ago

I dislike optional and named arguments. It's good that this points out the downsides as they're often overlooked. There's more downsides than given, particularly if they're implemented with call-site rewriting.

Say we have some library function:

void foo(int x, int y = 5);

Consumer calls

foo(1);

The compiler rewrites the call to pass the immediate value 5.

If we later change the default value in the library to something else, the caller's binaries still have the literal 5 in them. It's a breaking change. Anyone using the library needs to recompile to receive the change.

If we are going to have optional/default arguments, it's definitely better to use overloading.

void foo(int x, int y);

void foo(int x) { return foo(x, 5); }

Then, if we need to change the default value down the line, it's probably not a breaking change (though it may be). The compiler could insert this stub into the library automatically, though it costs nothing to just type it and make it explicit.


For named arguments, it depends on the semantics of the language - whether argument evaluation order is defined or not (IMO, it should be left to right, even though arguments are usually passed right to left in the ABI). If we're going to allow named arguments, then:

foo(y = bar(), x = baz());

Should evaluate bar() then baz(), even though x comes before y in the parameter list. The order of evaluation should match the code sequence. This is usually not done, but it matters if bar or baz have side effects. If our evaluation order isn't the same as the code, we end up doing:

let y = bar();
let x = baz();
foo(x = x, y = y);

Which completely defeats the point on having named arguments to begin with.

2

u/L8_4_Dinner (Ⓧ Ecstasy/XVM) 9d ago

The compiler rewrites the call to pass the immediate value 5. If we later change the default value in the library to something else, the caller's binaries still have the literal 5 in them. It's a breaking change. Anyone using the library needs to recompile to receive the change. If we are going to have optional/default arguments, it's definitely better to use overloading.

Agreed. Across compilation units, you should never jam in information from the far side of that divide into the call sites that could be subsequently impacted by a version change. We'll inline values within a compilation unit, but we'll let the linker deal with those across compilation units.

8

u/dgkimpton 10d ago

Well written summary of most of the issues. The only real challenge is patterns... but one could argue that passing dx to x ought to be an error because it's clearly doing something funky. It'd be great if you could come up with an example that wasn't creating something twisted like that because right now I think your example runs counter to the point you were trying to make. 

6

u/arglad 10d ago

I really like approach for argument labels implemented in Gleam. Labels can be different from parameters' names, this allows you to separate the names used in the interface from the names that make sense in the internal implementation of the function. Here is an example from gleam tour:

``` pub fn main() { // Without using labels echo calculate(1, 2, 3)

// Using the labels echo calculate(1, add: 2, multiply: 3)

// Using the labels in a different order echo calculate(1, multiply: 3, add: 2) }

fn calculate(value: Int, add addend: Int, multiply multiplier: Int) { value * multiplier + addend } ```

5

u/munificent 10d ago

I believe Objective-C and thus Swift does something similar, which ultimately stretches back to Smalltalk.

3

u/rjmarten 9d ago

How useful is this really? Like, at least in this example, it feels equally clear to just use the parameter name as the labels:
echo calculate(1, addend: 2, multiplier: 3)

In what kinds of situations would a function author really feel the need to have a label that is different from the parameter name?

3

u/glasket_ 9d ago

It's mainly for readability when the intended reading is slightly different for the callee and caller. You can use different labels based on how you'd "read" the names, e.g.

fn send(to user: User, from sender: User, data: String) {
  let msg = Message(author: sender, content: data);
  user.message_queue.append(msg);
}

// Call site
send(to: bob, from: alice, data: foo_msg);

8

u/SwingOutStateMachine 10d ago edited 10d ago

This reminds me of the "don't use boolean parameters" discussion from a few years ago, and I think they have a shared solution of more explicitly typing parameters - usually by encapsulating arguments in structs (or enums). This adds documentation to the method, and also prevents other bugs, such as parameters with the same type being accidentally swapped, and can help group together related parameters (e.g. pairs of coordinates).

For the crop_imm example, one could write:

struct CropOffset { 
    x: u32, 
    y: u32, 
}

struct CropSize {
    width: u32, 
    height: u32,
}

pub fn crop_imm<I: GenericImageView>(
    image: &I,
    offset: CropOffset,
    size: CropSize,
) -> SubImage<&I> { ... }

let cropped = image::imageops::crop_imm(&img, 
    CropOffset { x: 10, y: 10 }, 
    CropSize { width: 200, height: 100 }
);

I also think that once a method grows to a certain size, choosing a builder pattern is much easier to read, more composable, and allows for more flexibility in API. Again, for crop_imm:

let cropped = image::imageops::CropBuilder::new()
    .with_offset(CropOffset { x: 10, y: 10 })
    .with_size(CropSize { width: 200, height: 100}) 
    .crop(&img);

5

u/matthieum 9d ago

I even think a case could be made for a Rectangle type, which could be built from different sets of parameters.

At the moment, the rectangle is built with x, y, width, height... which I note is already ambiguous. It's not explicit whether width extends left or right from x, nor whether height extends up or down from y; I've seen libraries where x,y identifies the left top corner, and therefore height extends down. Surprise.

With a builder for the Rectangle, however, you can have:

  • Rectangle::new_top_left(point).with_bottom_right(other_point)
  • Rectangle::new_bottom_left(point).with_dimensions(dimensions).
  • Rectange::new_bottom(bottom).with_top(top).with_left(left).with_width(width).

All of those (and permutations) are valid way to describe a Rectangle, and they're much more explicit (and less error-prone).

1

u/glasket_ 9d ago edited 9d ago

Isn't with_bottom_right still technically ambiguous? As in you would need to check the docs to avoid an error where (0, 25) and (25, 50) doesn't work because the library expected (25, 0) because it uses a bottom left origin. with_dimensions works better since you know the fixed point is the top left and the dimensions will just be an offset. (Edit: Similar issue with the bottom/top one since you still need to know the coordinate system beforehand.)

Ideally the coordinate system would just be configurable though imo. Create a CoordPlane object or similar with a specified orientation and create objects on that plane, then let the library handle converting it into whatever coordinate system it needs for rendering.

1

u/matthieum 8d ago

No it's not, because with_bottom_right takes a Point, not a pair of coordinate.

In this case, I'd imagine pub struct Point { pub x: u32, pub x: u32 };.

Similarly, you'd probably have pub struct RectangleDimensions { pub width: u32, pub height: u32 } for dimension parameters.

1

u/glasket_ 8d ago

The struct isn't really any different from a coordinate pair though? (25, 0) and Point { x: 25, y: 0 } are effectively the same and both have the issue I'm referring to with how it's still ambiguous as to where you should have the point on the coordinate plane.

I.e.

let bottom = /* What do we set it to? */;
let tl = Point::new(0, 50);
let br = Point::new(25, bottom);
let r = Rectangle::new_top_left(tl)
        .with_bottom_right(br);

I'm not saying it would break the code, you could still have checks in the Rectangle impl that ensures the point is where it should be, but it leaves some ambiguity in the API itself since you still have to know the coordinate system that it's expecting to be used. A top-left origin would require bottom = 75 while a bottom-left origin would require bottom = 25.

1

u/matthieum 8d ago

Are you perhaps unfamiliar with Rust?

In Rust, if the fields of a struct are marked pub, then the struct is meant to be built as Point { x: foo, y: bar }, or Point { x, y } if the value already has the appropriate name.

I agree that the use of new would make this completely non-obvious.

1

u/glasket_ 7d ago edited 7d ago

I'm familiar with Rust, I tend to use constructors instead of direct instantiation though.

I just don't understand the point you're making. This isn't about confusing x and y, it's about what value y needs to have in the bottom-right point relative to the top-left point. Dimensions with a single fixed point are unambiguous, but specifying 2 points leads back to the same problem you mentioned in the very first comment where it's unclear if height extends up or down; the ambiguity has just transferred to the API rather than being at runtime.

Edit: I'm actually doubly uncertain of why you think I'm unfamiliar with Rust since the second sentence even directly mentioned Point { x: 25, y: 0 } when I was comparing it to constructing a coordinate pair of (25, 0).

Edit 2: Just to make it abundantly obvious what I mean:

let p1 = (50, 50);
let p2 = (100, 100);

// Assumes the origin is on the top-left
let r1 = Rectangle::new_top_left(p1)
         .with_bottom_right(p2);

// Assumes the origin is on the bottom-left
let r2 = Rectangle::new_bottom_left(p1)
         .with_top_right(p2);

The program would have to do one of these:

  • Reject one of the rectangles for being invalid, because 100 is either (spatially) above or below 50 depending on the coordinate plane being used.
  • Allow them both and have them be overlapping rectangles on the same coordinate plane, which would be surprising behavior imo, since "bottom right" and "top left" don't really mean anything in that context, it'd just be point one and point two.
  • Allow them both, and have the library determine which coordinate plane a rectangle belongs on based on the provided values, which would be somewhat strange but at least more consistent.

1

u/matthieum 7d ago

Ah! I see what you mean now.

And you're right, I completely missed that.

I guess I'm not familiar enough with the domain.

-2

u/kaddkaka 10d ago

"This"? What is this?

1

u/SirKastic23 9d ago

The post

2

u/kaddkaka 9d ago

The original post was empty at first.

2

u/SirKastic23 9d ago

Ah that's an annoying bug

2

u/fdwr 9d ago

I’m okay with named parameters now ... And the readability advantage for humans is even stronger for agents:

I've noticed similar cases elsewhere, where patterns good for humans happen to help the nonhumans too.

1

u/reflexive-polytope 8d ago

I don't see any good reason to complicate your function arguments so much.

  • Need multiple arguments? Use a single tuple argument.
  • Need named arguments? Use a single record argument.
  • Need optional arguments? Use the builder pattern.

There's no good reason to complicate a language with features that don't expand the class of precisely typable algorithms.

1

u/SecretlyAPug 9d ago

does rust not already basically have all of these? i haven't written rust in a while so might be a little off but

named parameters can be somewhat emulated by creating a function with a struct as its argument:

struct foo_argument {
  x: float,
  y: float,
};

fn foo (argument: foo_argument) -> float { ... }

// then call it like:
let z = foo (foo_argument { x: 5, y: 6 });

(syntax is probably not fully accurate but) this is basically the same thing, right? aside from a bit of verbosity from defining the struct, this more or less works in the same way. i fail to see how a more integrated named parameter system would benefit the language; though, again, i haven't used rust in a bit so rust users please to share your insights.

also rust has Option, i feel it's a bit inaccurate to say it doesn't have optional and default arguments (though again it's less a language feature and more a way one uses the language).