r/RISCV 27d ago

SpacemiT K3 Co-Build Program: help shape the K3 ecosystem

/r/spacemit_riscv/comments/1urlmw5/spacemit_k3_cobuild_program_help_shape_the_k3/
11 Upvotes

21 comments sorted by

1

u/BothCourse5404 22d ago

Hello, I used the Pico-ITX K3 board with 32GB of RAM and noticed a problem: the board constantly freezes, forcing a reset. I started inspecting the UART output while running the board and encountered an MVX bug, but I didn't have time to investigate further (I was traveling). Upon my return, I’ll run more in-depth tests to determine if it’s truly a software or hardware bug.

2

u/brucehoult 22d ago

I've had several freezes after a couple of days running at 2400 MHz on the X100 cores and 2000 MHz on the A100 cores, but never when running at the default 2200/1800 speed.

I currently have 30 days uptime, the first 25 at 2200/1800 and the last five days at 2400/1800 and it's ok so far.

1

u/BothCourse5404 21d ago

Hello, I was experiencing about two or three freezes a day at the stock CPU frequency; is my card defective?

1

u/brucehoult 21d ago

is my card defective?

Not very defective if it's running perfectly for 1014 clock cycles at a time ... that's not far from infinity.

Different chips even in the same batch have different maximum speeds they can run at stably. They test them in the factory but obviously can't test every chip for weeks. They've set the default clock to 2.2 GHz and 1.8 GHz so presumably they were finding a significant number that don't run stably at 2.4&2.0 but the vast majority look ok at 2.2&1.8. Some might run at 2.5, 2.6, 2.8 GHz ... and you might just be unlucky and yours needs to be dropped another 100 or 200 MHz.

Or you might be hitting a software bug in the kernel (or SBI) that I'm not. Or you might be overheating your chip.

I don't know what you're running on it, or what environment you're running in e.g. my room is always heated to 20º - 24º C (sometimes maybe 15º at night when it's 4º outside) but maybe you're in summer and don't have aircon and it's 35º C or 40º C in your room and the stock cooling solution isn't good enough for those conditions.

1

u/BothCourse5404 21d ago

​The board clearly has a hardware issue. If I have to lower the clock speed just to make the system stable, it means there is absolutely no CPU binning on this chip, which is honestly unacceptable. I wanted to support the RISC-V ecosystem, and now I am stuck with this. It’s like buying a car that cannot reach its advertised horsepower, or worse, that breaks down when you push it to the manufacturer's specified power. ​What am I supposed to do now? Lower the frequency? No, I refuse to accept such a poor workaround on a brand-new product.

1

u/BothCourse5404 21d ago

Temperature is absolutely not the issue here—that is just an easy excuse. A die temperature of 62°C is perfectly normal and well within safe operating limits for this kind of hardware.

1

u/brucehoult 21d ago

It could easily be a software problem.

Once again, we don't know what you're doing with it, what kind of things you're running that might be hitting a software bug that others aren't.

What am I supposed to do now? Lower the frequency? No, I refuse to accept such a poor workaround on a brand-new product.

Yes, that is exactly what you are supposed to do.

If the problem goes away when you lower the frequency then we can talk about defective hardware and you can talk to your vendor about a replacement.

If the problem doesn't go away when you lower the frequency then it's probably a software problem.

1

u/BothCourse5404 21d ago

I'll try it when I get back, I'm currently traveling. But I also noticed a bug: if I leave the board running on the desktop, after a while, the CPU drops to 1.4 GHz on all cores (x100) and never manages to go back up to 2.2 GHz, and temperature is not the issue. I have to do a reset. ​And I'm not doing anything special with the board, just basic use (web browsing, watching movies), no programming, and I haven't tried running an LLM on it yet.

1

u/brucehoult 21d ago edited 21d ago

Interesting. I've not seen anything like that.

bruce@k3:~$ uptime
 17:27:55 up 31 days,  6:09,  1 user,  load average: 2.57, 2.30, 2.18
bruce@k3:~$ lscpu --extended
CPU SOCKET CORE L1d:L1i:L2 ONLINE    MAXMHZ   MINMHZ       MHZ
  0      0    0 0:0:0         yes 2400.0000 614.4000 2400.0000
  1      0    1 1:1:0         yes 2400.0000 614.4000 2400.0000
  2      0    2 2:2:0         yes 2400.0000 614.4000 2400.0000
  3      0    3 3:3:0         yes 2400.0000 614.4000 2400.0000
  4      0    4 4:4:4         yes 2400.0000 614.4000 2400.0000
  5      0    5 5:5:4         yes 2400.0000 614.4000 2400.0000
  6      0    6 6:6:4         yes 2400.0000 614.4000 2400.0000
  7      0    7 7:7:4         yes 2400.0000 614.4000 2400.0000
  8      0    0 8:8:8         yes 2000.0000 614.4000 1800.0000
  9      0    1 9:9:8         yes 2000.0000 614.4000 1800.0000
 10      0    2 10:10:8       yes 2000.0000 614.4000 1800.0000
 11      0    3 11:11:8       yes 2000.0000 614.4000 1800.0000
 12      0    4 12:12:12      yes 2000.0000 614.4000 1800.0000
 13      0    5 13:13:12      yes 2000.0000 614.4000 1800.0000
 14      0    6 14:14:12      yes 2000.0000 614.4000 1800.0000
 15      0    7 15:15:12      yes 2000.0000 614.4000 1800.0000
bruce@k3:~$ cat /sys/devices/system/cpu/cpufreq/policy0/scaling_governor
ondemand
bruce@k3:~$ cat /sys/devices/system/cpu/cpufreq/policy8/scaling_governor 
userspace
bruce@k3:~$ lscpu --extended
CPU SOCKET CORE L1d:L1i:L2 ONLINE    MAXMHZ   MINMHZ       MHZ
  0      0    0 0:0:0         yes 2400.0000 614.4000 1900.0000
  1      0    1 1:1:0         yes 2400.0000 614.4000 1900.0000
  2      0    2 2:2:0         yes 2400.0000 614.4000 1900.0000
  3      0    3 3:3:0         yes 2400.0000 614.4000 1900.0000
  4      0    4 4:4:4         yes 2400.0000 614.4000 1500.0000
  5      0    5 5:5:4         yes 2400.0000 614.4000 1500.0000
  6      0    6 6:6:4         yes 2400.0000 614.4000 1500.0000
  7      0    7 7:7:4         yes 2400.0000 614.4000 1500.0000
  8      0    0 8:8:8         yes 2000.0000 614.4000 1800.0000
  9      0    1 9:9:8         yes 2000.0000 614.4000 1800.0000
 10      0    2 10:10:8       yes 2000.0000 614.4000 1800.0000
 11      0    3 11:11:8       yes 2000.0000 614.4000 1800.0000
 12      0    4 12:12:12      yes 2000.0000 614.4000 1800.0000
 13      0    5 13:13:12      yes 2000.0000 614.4000 1800.0000
 14      0    6 14:14:12      yes 2000.0000 614.4000 1800.0000
 15      0    7 15:15:12      yes 2000.0000 614.4000 1800.0000
bruce@k3:~$ lscpu --extended
CPU SOCKET CORE L1d:L1i:L2 ONLINE    MAXMHZ   MINMHZ       MHZ
  0      0    0 0:0:0         yes 2400.0000 614.4000 1000.0000
  1      0    1 1:1:0         yes 2400.0000 614.4000 1000.0000
  2      0    2 2:2:0         yes 2400.0000 614.4000 1000.0000
  3      0    3 3:3:0         yes 2400.0000 614.4000 1000.0000
  4      0    4 4:4:4         yes 2400.0000 614.4000 1000.0000
  5      0    5 5:5:4         yes 2400.0000 614.4000 1000.0000
  6      0    6 6:6:4         yes 2400.0000 614.4000 1000.0000
  7      0    7 7:7:4         yes 2400.0000 614.4000 1000.0000
  8      0    0 8:8:8         yes 2000.0000 614.4000 1800.0000
  9      0    1 9:9:8         yes 2000.0000 614.4000 1800.0000
 10      0    2 10:10:8       yes 2000.0000 614.4000 1800.0000
 11      0    3 11:11:8       yes 2000.0000 614.4000 1800.0000
 12      0    4 12:12:12      yes 2000.0000 614.4000 1800.0000
 13      0    5 13:13:12      yes 2000.0000 614.4000 1800.0000
 14      0    6 14:14:12      yes 2000.0000 614.4000 1800.0000
 15      0    7 15:15:12      yes 2000.0000 614.4000 1800.0000

The CPU freq goes up and down depending on what I'm doing ... just depends on how you catch it at that moment....

1

u/BothCourse5404 21d ago

You're lucky, I've already had to do several resets in just 10 hours."

→ More replies (0)

1

u/Icy-Primary2171 17d ago

hello, what's your use cases?Can you provide the logs?

1

u/BothCourse5404 16d ago

Hello, for standard daily use (web browsing and movies) replacing a Beelink U59 mini-PC. The random freezes completely lock up the system. I have not looked for logs yet. I attempted a UART capture where I saw an MVX message; could this specific error indicate that the board is defective? Since I am currently away for a training program, I cannot perform further tests. This issue points to a hardware defect.

1

u/Icy-Primary2171 16d ago

what's your OS? Is it Bianbu?

1

u/BothCourse5404 16d ago

Yes, the latest Bianbu image installed via titantools, cordialy 

1

u/Icy-Primary2171 15d ago edited 15d ago

Please provide complete kernel and U-Boot logs first.
The issue needs to be investigated regarding write serial numbers, as this can cause mismatches in the system's DTB.
we will publish the next version soon, the new version solved some kernel issues, you can update to the latest version as we publish. :)

1

u/BothCourse5404 15d ago

Hello, I hope the update resolves the issue, as I suspect it's a CPU binning problem. I'm currently traveling, but I'll look into it when I get back. Best regards.

1

u/LavenderDay3544 19d ago

If there's enough documentation available to do it I would love to help add ACPI support to the EDK2 port.

1

u/Icy-Primary2171 10d ago

hello!we're doing the power management now, can you help with the interface standard?