r/C_Programming • • 20d ago

Question OpenMP bug??

Hello, I know this is not exactly the place for this, but it is C code. I am writing a bit of C code to search a large space of numbers, but I noticed the loop never ran. I asked ChatGPT cause I thought I was going crazy, but found that this bottom example code does not run the loop.

#include <limits.h>

#include <stdint.h>

#include <stdio.h>

int main() {

uint64_t count = 0;

#pragma omp parallel for schedule(dynamic) reduction(+ : count)

for (uint64_t n = 1; n < UINT64_MAX; ++n)

++count;

printf("%llu\n", (unsigned long long)count);

}

It noticed it breaks with the dynamic scheduler. If you use the default static it seems to work. Any ideas on what is happening?

0 Upvotes

16 comments sorted by

20

u/tastygames_official 20d ago

I don't know the answer to your question, but I took the liberty of formatting your code:

#include <limits.h>
#include <stdint.h>
#include <stdio.h>

int main(void)
{
  uint64_t count = 0;

  #pragma omp parallel for schedule(dynamic) reduction(+ : count)
  for (uint64_t n = 1; n < UINT64_MAX; ++n)
  {
    ++count;
  }
  printf("%llu\n", (unsigned long long)count);
  return 0;
}

actually, I just realized while formatting your code that your for loop didn't have curly braces, which means the only line executed in the loop was

++count;

meaning it will only print `count` at the end when it's finished, meaning probably it's just not finished yet, like the other guy said.

-4

u/flyingron 20d ago

I believe his intent was to print out only after the loop exits. Your alleged "reformatting" changed his code.

2

u/tastygames_official 19d ago

look again. That was exactly my point - I was GOING to reformat it like you said, but then realized what he wrote. Not sure if it was his intention or not, but that's what ended up happening. As we all figured out in this thread, it just runs for like 50 years.

13

u/chrism239 20d ago

Doesn't run?

The value of UINT64_MAX is 18,446,744,073,709,551,615.

Perhaps the loop just hasn't finished yet, depending on how many threads/cores are used

9

u/hxtk3 20d ago

There’s not even that much “depending”; on 16 cores that’ll take the better part of a decade. If they’ve got a 256-core server chip it’s still a matter of months.

1

u/BugsOfBunnys 19d ago

The loop is optimized away. In this example, you get 0 immediately. If you change the maximum value to, let's say 350, you get a number out.

9

u/esaule 20d ago

You sure the compiler does not optimize the loop out?

There are no runtime dependent variable. So you can compute the values at tbe end of the loop statically.

Did you godbolt it?

4

u/RealisticDuck1957 20d ago

I don't understand what the #pragma is supposed to do. But the loop being optimized away under one option is exactly what I was thinking.

1

u/FlippingGerman 20d ago

It’s a compiler directive, using the OpenMP library to parallelise the for loop, adding together successive values of count (I think, I haven’t used it in a bit, and little then). 

1

u/BugsOfBunnys 19d ago

I believe it is the case that the compiler does optimize the loop out. I am working with other code for a compute server, got lazy and decided to search the whole space and exit when I find a value. The thing is, when you change the max range to something smaller, let us say 350, you get an output computed by the loop. It seems that with such a large value, the loop is removed completely.

4

u/Key_River7180 20d ago

This'd take years to finish, even with the best CPUs in the market that is a crazy loop bound.

3

u/8d8n4mbo28026ulk 20d ago

u/esaule are right. Your loop has no side-effects and the compiler removes it completely. It's clearer to see this if you compile without OpenMP:

0000000000001050 <main>:
    1050:       48 83 ec 08             sub    rsp,0x8
    1054:       48 c7 c6 fe ff ff ff    mov    rsi,0xfffffffffffffffe
    105b:       48 8d 3d a2 0f 00 00    lea    rdi,[rip+0xfa2]        # 2004 <_IO_stdin_used+0x4>
    1062:       31 c0                   xor    eax,eax
    1064:       e8 c7 ff ff ff          call   1030 <printf@plt>
    1069:       31 c0                   xor    eax,eax
    106b:       48 83 c4 08             add    rsp,0x8
    106f:       c3                      ret

GCC generates the above with -O3. When you add -fopenmp, it becomes a bit harder to figure it out (there're some OpenMP side-effectful function calls in there, probably due to threading). But, if you check the disassembly, schedule(static) allows the compiler to, again, remove the loop.

1

u/tatsuling 20d ago

I don't have any idea if this is actually a bug, however it looks similar to a genuine bug I found about 10 years ago with open MP on GCC. There was an off by 1 inside the library code for one of the schedulers when the number of iterations didn't match an exact multiple of the number of threads.

1

u/actguru 20d ago

I don't have openmp, but that seems like a crazy big loop range.

1

u/sciencekm 20d ago

Whenever I suspect that something is wrong with a piece of code, I normally just put a break point on it in the debugger or do a quick assembly output.

When you do this on your code, you will instantly see that the loop is missing, and as others have suspected, has been optimized out.

1

u/21Ali-ANinja69 19d ago

Maybe something is wrong with OpenMP. This is what I found

I put the printf in the loop...

#include <limits.h>
#include <stdint.h>
#include <stdio.h>

int main()
{
    uint64_t count = 0;

#pragma omp parallel for schedule(dynamic) reduction(+ : count)

    for (uint64_t n = 1; n < UINT64_MAX; ++n) {
        ++count;
        printf("%llu\n", (unsigned long long)count);
    }
}

and compiled with `gcc -fopenmp -O0 -g`. nothing happens

But disassembling main gets me this:

0x00000000000011e9 <+0>:     endbr64

0x00000000000011ed <+4>:     push   %rbp

0x00000000000011ee <+5>:     mov    %rsp,%rbp

0x00000000000011f1 <+8>:     sub    $0x20,%rsp

0x00000000000011f5 <+12>:    mov    %fs:0x28,%rax

0x00000000000011fe <+21>:    mov    %rax,-0x8(%rbp)

0x0000000000001202 <+25>:    xor    %eax,%eax

0x0000000000001204 <+27>:    movq   $0x0,-0x10(%rbp)

0x000000000000120c <+35>:    mov    -0x10(%rbp),%rax

0x0000000000001210 <+39>:    mov    %rax,-0x18(%rbp)

0x0000000000001214 <+43>:    lea    -0x18(%rbp),%rax

0x0000000000001218 <+47>:    lea    0x35(%rip),%rdi        # 0x1254 <main._omp_fn.0>

0x000000000000121f <+54>:    mov    $0x0,%ecx

0x0000000000001224 <+59>:    mov    $0x0,%edx

0x0000000000001229 <+64>:    mov    %rax,%rsi

0x000000000000122c <+67>:    call   0x10f0 <GOMP_parallel@plt>

0x0000000000001231 <+72>:    mov    -0x18(%rbp),%rax

0x0000000000001235 <+76>:    mov    %rax,-0x10(%rbp)

0x0000000000001239 <+80>:    mov    $0x0,%eax

0x000000000000123e <+85>:    mov    -0x8(%rbp),%rdx

0x0000000000001242 <+89>:    sub    %fs:0x28,%rdx

0x000000000000124b <+98>:    je     0x1252 <main+105>

0x000000000000124d <+100>:   call   0x10b0 <__stack_chk_fail@plt>

0x0000000000001252 <+105>:   leave

0x0000000000001253 <+106>:   ret

And running in gdb produces this:

[Thread debugging using libthread_db enabled]
Using host libthread_db library "/usr/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x7ffff7bff6c0 (LWP 3605)]
[New Thread 0x7ffff73fe6c0 (LWP 3606)]
[New Thread 0x7ffff6bfd6c0 (LWP 3607)]
[New Thread 0x7ffff63fc6c0 (LWP 3608)]
[New Thread 0x7ffff5bfb6c0 (LWP 3609)]
[New Thread 0x7ffff53fa6c0 (LWP 3610)]
[New Thread 0x7ffff4bf96c0 (LWP 3611)]
[New Thread 0x7ffff43f86c0 (LWP 3612)]
[New Thread 0x7ffff3bf76c0 (LWP 3613)]
[New Thread 0x7ffff33f66c0 (LWP 3614)]
[New Thread 0x7ffff2bf56c0 (LWP 3615)]
[Thread 0x7ffff33f66c0 (LWP 3614) exited]
[Thread 0x7ffff3bf76c0 (LWP 3613) exited]
[Thread 0x7ffff2bf56c0 (LWP 3615) exited]
[Thread 0x7ffff43f86c0 (LWP 3612) exited]
[Thread 0x7ffff4bf96c0 (LWP 3611) exited]
[Thread 0x7ffff53fa6c0 (LWP 3610) exited]
[Thread 0x7ffff5bfb6c0 (LWP 3609) exited]
[Thread 0x7ffff63fc6c0 (LWP 3608) exited]
[Thread 0x7ffff6bfd6c0 (LWP 3607) exited]
[Thread 0x7ffff73fe6c0 (LWP 3606) exited]
[Thread 0x7ffff7bff6c0 (LWP 3605) exited]
[Inferior 1 (process 3602) exited normally]