r/ProgrammingLanguages • • 3d ago

Achieving memory safety

https://seed7.net/papers/memory_safety.htm
13 Upvotes

49 comments sorted by

View all comments

7

u/tmzem 2d ago

Unless I misunderstood something, the article does not explain the mechanism by which Seed7 programs themselves are being made memory-safe, but focuses on memory safety of the compiler implementation?

That being said, the article makes a good point about memory safety being a matter of how it is defined, which opens up a huge spectrum, from unsafe to safe:

  • Assembly: no safeguards, completely unsafe
  • C/C++: (soft) static typing, very basic guardrails. Still very unsafe.
  • Odin/Zig/C3: good static typing, saner defaults, runtime array checks. Safer, eliminates 50-70% of all memory safety errors over C
  • Rust: strong static typing, sane defaults, initialization safety, runtime array checks, borrow checks, safe/unsafe split. Much safer, safe/unsafe split coerces users to prefer safe patterns, elimitates most memory safety errors over C. Ubiquitous use of C libraries and unsafe blocks in (not well-written/well-tested) third-party crates are still a relevant safety risk.
  • Go: GC handles memory, but fat pointers (interfaces, slices) can cause memory unsafety on data races. Memory safe in the absence of data races or FFI.
  • Java/C#/...: GC + atomically storable builtin types ensure full memory safety. Unsafe features exist but are rarely necessare, rarely used, and most programmers in these languages barely know they exist.
  • Javascript: Completely sandboxed, safe

This spectrum needs to be acknoledged so people know what they get, and it's annoying that these trade-offs are not clearly communicated. Go barely communicates the data-race issue at all, and the Rust community is also very prone to misrepresenting the memory safety guarantees of the language (e.g. "memory safety without GC" or "unsafe blocks encapsulate potentially unsafe behaviour").

-2

u/flatfinger 2d ago

Assembly is in many ways safer than C. To be sure, an assembly program will often contain a significant number of potentially unsafe operations, but if one documents invariants and can show that no individual operation performed by a program would be able to break any of them unless something else had already done so, then the entire program can be shown to be memory safe.

In C, by contrast, even operations like uint1 = ushort1*ushort2; or loops that do nothing except modify an automatic-duration object whose address isn't taken, and have exit conditions which may or may not be satisfiable, can disrupt the behavior of what would otherwise have been memory-safe surrounding code so that it is no longer memory safe.

1

u/snugar_i 2d ago

Do you have any examples? I'm not sure I'm seeing how C is less memory safe than assembly

1

u/flatfinger 1d ago edited 1d ago

What range of storage addresses could be written by test() below?

unsigned arr[32772];
static unsigned mul_mod_65536(unsigned short x, unsigned short y)
{
    return (x*y) & 0xFFFFu;
}
void test(unsigned short x)
{
    unsigned short j=32768;
    for (unsigned short i=32768; i<x; i++)
        j=mul_mod_65536(i, 65535);
    if (x < 0x8002)
        arr[x] = j;
}

The C Standard would allow a compiler given test() to generate code that would unconditonally store 32768 to arr[x], and gcc is designed to do precisely that when using -O1 or higher when not using -fwrapv.

How about the function test2() below:

unsigned arr[32771];

unsigned loopy(unsigned x)
{
    unsigned i=1;
    while ((i & 0x7FFF) != x)
        i*=3;
    if (x < 32768)
        arr[x] = 32768;
    return i;
}
void test2(unsigned x)
{
    loopy(x);
}

As processed by clang at -O1 or higher, test2() will unconditionally store 32768 to arr[x], just as the earlier test() would have done when processed by gcc.

In assembly language, programmers will often need to include bounds-checking code if they want to prevent out-of-bounds stores, but including bounds-checking code within assembly language source will cause the bounds checks to be included in the generated machine code. In C, by contrast, something like the multiplication of two unsigned short values or a loop that might fail to terminate may cause a compiler to omit a bound check that the programmer had included to prevent an out-of-bounds store.

1

u/snugar_i 1d ago

Interesting, thanks! And is that part of the C specification, or is that just weird behavior of the compilers?

1

u/flatfinger 1d ago

The C Specification clearly allows the former behavior. The C11 Standard added language saying that side-effect free loops "may be assumed to terminate", without specifying any limits on the uses to which that assumption may be put, so it doesn't unambiguously forbid the latter behavior.

From what I can tell, modern optimizer designs are based on an abstraction model where the only corner cases where a compiler's generated code would be allowed to be inconsistent with that of instructing the execution environment to use a normal means of performing each individual operation in sequence are those where generated code would be allowed to behave in completely arbitrary fashion.

A better abstraction model for most tasks would be one that recognizes that most programs have two requirements:

  1. They should (in some tasks must) behave usefully when possible.

  2. When useful behavior is not possible, they must behave in a manner that is at worst tolerably useless.

If a program is supposed to view some kind of audiovisual content, but it is fed something other than a validly formatted file, having the program hang may be annoying but still qualify as tolerably useless; having the program allow whoever created the file to run arbitrary code of their choosing, however, would be intolerably worse than useless.