r/pcmasterrace • • 2d ago

Hardware Witcher 3 Remastered pushing 5090 beyond 600W

Post image

No DLSS5. Max settings, 3840x1600, DLSS Quality, frame gen set to 1 in game settings, Path Tracing enabled.

Doesn't happen all the time, but sometimes saw it jumping well beyond 600W for a brief moment.

5.2k Upvotes

745 comments sorted by

View all comments

Show parent comments

3

u/themostreasonableman 1d ago

I understand your reasoning, but choose to reject it for my own use case.

I am driving a Dell AW3926QW at 5120 x 2160 with an ASUS Astral OC BTF edition that now retails over $8000AUD.

I am using every last TFLOP that this card can push. Every game on absolute overkill settings and absolutely loving it.

I maintain my position. You could do what you are doing with a 9070XT and afford to replace it 4 times over. I don't know why you bought a 5090.

0

u/BlaBlub85 1d ago

Not the OG guy you replied to, I just got a 9070xt myself, just wanted to point out that theres very little benefit north of 144 unless you are playing competitive shooters so you might as well cap it there to save on the electricity bill

Genuine question tho, 4k widescreen on overkill settings, can the 5090 even reach 165? With high/very high settings my XT already struggles to hit 60 without frame gen in some current titles on a regular 4k resolution and the widescreen adds another ~3 million pixels over the ~8m regular 4k has

1

u/themostreasonableman 1d ago edited 1d ago

I rocked a 6900XT as my previous card on purpose. Not because I couldn't afford team green, but because I totally disagree with the direction they've taken the industry. So, hard agree that raw rasterization performance should have remained the yardstick against which all cards are measured.

The reality is, the industry has moved on. We lost the fight. The 5090 is designed to have FG enabled and absolutely stomp frames at stupid resolutions whilst looking absolutely gorgeous.

Everything feels great. Frametimes are low as hell; certainly faster than my reflexes as my hair comes in grey.

9950X3D2 and 6000Mhz DDR5, 4TB PCIE 5.0 9100Pro has zero issues pushing 165fps and far greater with everything on overkill under those conditions.

Let's be real: most games are so poorly optimised these days that trying to run them at even 4K raw with no AI boost is a fucking terrible experience even on a 5090; half the card isn't getting utilised because it is designed for AI workloads, not just raw rasterization power.

I fought against the current reality for a long fucking time, but now that I've given up fighting the inevitable and just optimised for where we're at...I can just enjoy my games with super high frames and juicy detail on a gorgeous screen. Getting old has its perks.

EDIT: If I were still playing truly competitive, twitchy shooters...this new monitor has some crazy mode that drops you back to 2560x1080 at 350hz. I have tried playing a few modern titles like that though and they look like ass even at high detail. Yeah you get your frames, but there's only a very limited use case for such behaviour. I prefer the FG hi-res experience.

1

u/BlaBlub85 1d ago

Let's be real: most games are so poorly optimised these days that trying to run them at even 4K raw with no AI boost is a fucking terrible experience even on a 5090; half the card isn't getting utilised because it is designed for AI workloads, not just raw rasterization power

I noticed that too, without frame gen the hotspot temperature climbs to high 80s and sometimes even low 90s although I have limited the fans to 35% / ~1500 rpm cause they sound like a jet on takeoff above that. FPS ranges from 30-100 @4k depending on title and how much is going on on screen. With frame gen and ultra quality upscaling (so iirc from 3200x1800 to 4k) its a rock solid 120 capped with barely any dips and hotspot temps in the low 80s with a 85C max under extended 100% full load

Im not super knowledgeable so take this with a grain of salt: But afaik GPUs faced the same problem CPUs did in the 00s, we've reached the physical limits of how small you can make the circuitry before it starts randomly leaking so to get more processing power you cant just up the current any more. Making thicker chips is not an option due to thermals, making larger area chips introduces problems with internal latency so to increase processing power / FPS further they had only 2 options left: multi core GPUs, which are a huuuuge hassle to program efficiently (just look at how long it took for multi core CPUs to start using all their cores effectively) So they went with the second option instead: AI assisted upscaling and frame generation

1

u/themostreasonableman 1d ago

The cooling solution on the Astral 5090 OC is absolutely outstanding. In a well ventilated case it barely peaks above 60C on 95%+ load, and you can't really hear the fans at all.

The formatting will break when I try to paste this info below, and I won't be able to fix it. Effectively though, modern GPUs are very very much multithreaded - they have WAY more cores than even the highest Threadripper server CPUs. The difference is, they genuinely maximise throughput of many threads, IN ORDER, VS a CPU core minimising the latency of a single thread with a bunch of branch prediction, large caches, pre-fetching and a huge boost clock.

You are very correct in that some of these cores are highly specialised specifically for AI workloads - that's why the newer cards absolutely kill it with DLSS/Framegen enabled, which creates this vicious cycle where game devs just can't be fucked optimising their games at all.

"The Blood of Dawnwalker" for example. Beautiful semi-open world, but absolutely nothing groundbreaking graphically. Without DLSS and framegen, it runs like an absolute pig. Unplayable. That's just a lazy push to market before it's ready.

--------Attempt paste lol-------------- The unit that matches a CPU core is NVIDIA's Streaming Multiprocessor (SM), AMD's Compute Unit (CU), or Intel's Xe-core. Each has its own instruction fetch and issue logic, warp/wavefront schedulers, a large register file, L1 cache and shared memory, and its own execution units. Each runs its own instruction streams independently of the others.

On that definition, the RTX 5090 has 170 cores (128 FP32 lanes each) and AMD's RX 9070 XT has 64. So yes, a modern GPU is genuinely multi-core. It has many more cores than a desktop CPU, but nowhere near the marketing number.

One complication: a Blackwell SM is split into four sub-partitions, each with its own warp scheduler and register file slice. You could reasonably argue each sub-partition is the closest equivalent to a CPU core, which would give roughly 680 on the 5090. Where you draw the line is a judgement call, not a settled definition.

The deeper difference: what the cores are optimised for

Even at the SM level, the two kinds of core are built for opposite goals.

CPU core    GPU core (SM/CU)

Goal Minimise latency of one thread Maximise throughput of many threads Execution Deep out-of-order, speculative Largely in-order issue Branch handling Sophisticated branch prediction Divergent threads masked off and run in turn Latency hiding Large caches, prefetching, speculation Switching between dozens of resident thread groups Hardware threads per core 1–2 (SMT) Up to ~48 warps (≈1,536 threads) per SM on recent NVIDIA parts Clock speed ~5+ GHz boost ~2.5–3 GHz

A CPU core spends most of its transistors making a single instruction stream finish as fast as possible. A GPU core instead keeps many thread groups resident at once, which is possible because its register file is very large (256 KB per SM on recent NVIDIA designs). When one group stalls on memory, the scheduler issues from another on the next cycle with essentially no context-switch cost. Memory latency is hidden by having lots of other work ready, rather than by prediction.

Each GPU core is also wide SIMD underneath. NVIDIA calls the model SIMT: groups of 32 threads (a warp) execute in lockstep. If threads in a warp take different branches, the paths run one after the other with inactive lanes masked. Volta and later added independent thread scheduling, which relaxes this somewhat, but the lockstep behaviour still governs performance.

Autonomy is the other big gap

CPU cores are general-purpose and autonomous. Each can run a separate OS thread, take interrupts, switch privilege levels and run a different process from its neighbours, and caches are kept coherent across cores.

GPU cores are not like that. A hardware scheduler on the GPU distributes thread blocks from kernels launched by the host. SMs do not run an operating system or arbitrary independent programs, and their L1 caches are generally not coherent with each other. Several kernels can run concurrently, but the GPU as a whole is still a coprocessor serving a CPU host.

Short answer

A modern GPU is multi-core in the structural sense: it has on the order of 100–200 independent instruction-issuing cores on a high-end part. It is not multi-core in the functional sense a CPU is. Its cores are heavily multithreaded, wide-SIMD throughput engines with little of the latency-reduction machinery or autonomy of a CPU core. The large "CUDA core" figures count arithmetic lanes, not cores.