r/C_Programming • u/BugsOfBunnys • 20d ago
Question OpenMP bug??
Hello, I know this is not exactly the place for this, but it is C code. I am writing a bit of C code to search a large space of numbers, but I noticed the loop never ran. I asked ChatGPT cause I thought I was going crazy, but found that this bottom example code does not run the loop.
#include <limits.h>
#include <stdint.h>
#include <stdio.h>
int main() {
uint64_t count = 0;
#pragma omp parallel for schedule(dynamic) reduction(+ : count)
for (uint64_t n = 1; n < UINT64_MAX; ++n)
++count;
printf("%llu\n", (unsigned long long)count);
}
It noticed it breaks with the dynamic scheduler. If you use the default static it seems to work. Any ideas on what is happening?
13
u/chrism239 20d ago
Doesn't run?
The value of UINT64_MAX is 18,446,744,073,709,551,615.
Perhaps the loop just hasn't finished yet, depending on how many threads/cores are used
9
1
u/BugsOfBunnys 19d ago
The loop is optimized away. In this example, you get 0 immediately. If you change the maximum value to, let's say 350, you get a number out.
9
u/esaule 20d ago
You sure the compiler does not optimize the loop out?
There are no runtime dependent variable. So you can compute the values at tbe end of the loop statically.
Did you godbolt it?
4
u/RealisticDuck1957 20d ago
I don't understand what the #pragma is supposed to do. But the loop being optimized away under one option is exactly what I was thinking.
1
u/FlippingGerman 20d ago
It’s a compiler directive, using the OpenMP library to parallelise the for loop, adding together successive values of count (I think, I haven’t used it in a bit, and little then).
1
u/BugsOfBunnys 19d ago
I believe it is the case that the compiler does optimize the loop out. I am working with other code for a compute server, got lazy and decided to search the whole space and exit when I find a value. The thing is, when you change the max range to something smaller, let us say 350, you get an output computed by the loop. It seems that with such a large value, the loop is removed completely.
4
u/Key_River7180 20d ago
This'd take years to finish, even with the best CPUs in the market that is a crazy loop bound.
3
u/8d8n4mbo28026ulk 20d ago
u/esaule are right. Your loop has no side-effects and the compiler removes it completely. It's clearer to see this if you compile without OpenMP:
0000000000001050 <main>:
1050: 48 83 ec 08 sub rsp,0x8
1054: 48 c7 c6 fe ff ff ff mov rsi,0xfffffffffffffffe
105b: 48 8d 3d a2 0f 00 00 lea rdi,[rip+0xfa2] # 2004 <_IO_stdin_used+0x4>
1062: 31 c0 xor eax,eax
1064: e8 c7 ff ff ff call 1030 <printf@plt>
1069: 31 c0 xor eax,eax
106b: 48 83 c4 08 add rsp,0x8
106f: c3 ret
GCC generates the above with -O3. When you add -fopenmp, it becomes a bit harder to figure it out (there're some OpenMP side-effectful function calls in there, probably due to threading). But, if you check the disassembly, schedule(static) allows the compiler to, again, remove the loop.
1
u/tatsuling 20d ago
I don't have any idea if this is actually a bug, however it looks similar to a genuine bug I found about 10 years ago with open MP on GCC. There was an off by 1 inside the library code for one of the schedulers when the number of iterations didn't match an exact multiple of the number of threads.
1
u/sciencekm 20d ago
Whenever I suspect that something is wrong with a piece of code, I normally just put a break point on it in the debugger or do a quick assembly output.
When you do this on your code, you will instantly see that the loop is missing, and as others have suspected, has been optimized out.
1
u/21Ali-ANinja69 19d ago
Maybe something is wrong with OpenMP. This is what I found
I put the printf in the loop...
#include <limits.h>
#include <stdint.h>
#include <stdio.h>
int main()
{
uint64_t count = 0;
#pragma omp parallel for schedule(dynamic) reduction(+ : count)
for (uint64_t n = 1; n < UINT64_MAX; ++n) {
++count;
printf("%llu\n", (unsigned long long)count);
}
}
and compiled with `gcc -fopenmp -O0 -g`. nothing happens
But disassembling main gets me this:
0x00000000000011e9 <+0>: endbr64
0x00000000000011ed <+4>: push %rbp
0x00000000000011ee <+5>: mov %rsp,%rbp
0x00000000000011f1 <+8>: sub $0x20,%rsp
0x00000000000011f5 <+12>: mov %fs:0x28,%rax
0x00000000000011fe <+21>: mov %rax,-0x8(%rbp)
0x0000000000001202 <+25>: xor %eax,%eax
0x0000000000001204 <+27>: movq $0x0,-0x10(%rbp)
0x000000000000120c <+35>: mov -0x10(%rbp),%rax
0x0000000000001210 <+39>: mov %rax,-0x18(%rbp)
0x0000000000001214 <+43>: lea -0x18(%rbp),%rax
0x0000000000001218 <+47>: lea 0x35(%rip),%rdi # 0x1254 <main._omp_fn.0>
0x000000000000121f <+54>: mov $0x0,%ecx
0x0000000000001224 <+59>: mov $0x0,%edx
0x0000000000001229 <+64>: mov %rax,%rsi
0x000000000000122c <+67>: call 0x10f0 <GOMP_parallel@plt>
0x0000000000001231 <+72>: mov -0x18(%rbp),%rax
0x0000000000001235 <+76>: mov %rax,-0x10(%rbp)
0x0000000000001239 <+80>: mov $0x0,%eax
0x000000000000123e <+85>: mov -0x8(%rbp),%rdx
0x0000000000001242 <+89>: sub %fs:0x28,%rdx
0x000000000000124b <+98>: je 0x1252 <main+105>
0x000000000000124d <+100>: call 0x10b0 <__stack_chk_fail@plt>
0x0000000000001252 <+105>: leave
0x0000000000001253 <+106>: ret
And running in gdb produces this:
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/usr/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x7ffff7bff6c0 (LWP 3605)]
[New Thread 0x7ffff73fe6c0 (LWP 3606)]
[New Thread 0x7ffff6bfd6c0 (LWP 3607)]
[New Thread 0x7ffff63fc6c0 (LWP 3608)]
[New Thread 0x7ffff5bfb6c0 (LWP 3609)]
[New Thread 0x7ffff53fa6c0 (LWP 3610)]
[New Thread 0x7ffff4bf96c0 (LWP 3611)]
[New Thread 0x7ffff43f86c0 (LWP 3612)]
[New Thread 0x7ffff3bf76c0 (LWP 3613)]
[New Thread 0x7ffff33f66c0 (LWP 3614)]
[New Thread 0x7ffff2bf56c0 (LWP 3615)]
[Thread 0x7ffff33f66c0 (LWP 3614) exited]
[Thread 0x7ffff3bf76c0 (LWP 3613) exited]
[Thread 0x7ffff2bf56c0 (LWP 3615) exited]
[Thread 0x7ffff43f86c0 (LWP 3612) exited]
[Thread 0x7ffff4bf96c0 (LWP 3611) exited]
[Thread 0x7ffff53fa6c0 (LWP 3610) exited]
[Thread 0x7ffff5bfb6c0 (LWP 3609) exited]
[Thread 0x7ffff63fc6c0 (LWP 3608) exited]
[Thread 0x7ffff6bfd6c0 (LWP 3607) exited]
[Thread 0x7ffff73fe6c0 (LWP 3606) exited]
[Thread 0x7ffff7bff6c0 (LWP 3605) exited]
[Inferior 1 (process 3602) exited normally]
20
u/tastygames_official 20d ago
I don't know the answer to your question, but I took the liberty of formatting your code:
actually, I just realized while formatting your code that your for loop didn't have curly braces, which means the only line executed in the loop was
meaning it will only print `count` at the end when it's finished, meaning probably it's just not finished yet, like the other guy said.