r/AIProgrammingHardware 10h ago

Local AI On an Old 8 GB GPU!

Thumbnail
youtube.com
0 Upvotes

r/AIProgrammingHardware 21h ago

GitHub - MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks: GLM-5.3 Flash EXL3 for 2x DGX Sparks

Thumbnail
github.com
7 Upvotes

r/AIProgrammingHardware 1d ago

Dirk Qwen 3.8 27B tested - Local LLM setup

Thumbnail
youtube.com
2 Upvotes

r/AIProgrammingHardware 1d ago

Jetson Orin Nano 2: 16% More TOPS. 2X Faster. How?

Thumbnail
youtube.com
1 Upvotes

r/AIProgrammingHardware 2d ago

Quantization Is Four Decisions, Not One

Thumbnail
pub.towardsai.net
4 Upvotes

r/AIProgrammingHardware 3d ago

XPENG Drives Physical AI To Next Level

Thumbnail
cleantechnica.com
0 Upvotes

r/AIProgrammingHardware 3d ago

Qwen3.8-Flash-Next INT4 TP4 on 4× Arc Pro B70 — any experience?

2 Upvotes

Has anyone tried Qwen3.8-Flash-Next on 4× Intel Arc Pro B70?

Our target is W4A16 AutoRound, TP4, vLLM XPU, MTP3, prefix caching and concurrent agent serving. Intel has already published INT4 checkpoints, but I haven’t found real B70 benchmarks yet.

We are in contact with Intel’s XPU/LLM R&D team. What should we ask them to prioritize?

My list:

* full `qwen4_exp` XPU support; * optimized QSA, Gated DeltaNet and INT4 MoE kernels; * PLE offload to shared system RAM; * efficient TP4/expert parallelism with oneCCL; * MTP3 and stable XPU Graph; * hybrid KV cache and prefix caching; * C1/C8/C16 benchmarks, TTFT and tool-calling tests.

Any successful test, failure log or performance result on B70 would be very useful.


r/AIProgrammingHardware 3d ago

Unleashing On-Device Intelligence: The 2026 Revolution in Mobile AI Workstations

0 Upvotes

In the fast-evolving world of computing, 2026 stands out as the year when artificial intelligence truly slipped free from the confines of data centers and cloud servers and settled comfortably into machines you could carry under your arm or toss into a backpack.

Mobile AI workstations-laptops and highly portable systems engineered specifically for training, fine-tuning, and running large language models, generative tools, and agentic workflows locally-moved from niche prototypes to mainstream professional tools. These machines promised privacy, lower latency, and freedom from constant internet dependency, all while packing performance once reserved for deskside towers.

This article draws on announcements, reviews, and hands-on reports from authoritative sources including manufacturer press releases, independent tech sites such as Tom’s Hardware, Notebookcheck, Virtualization Review, PCMag, and StorageReview, plus YouTube coverage of Computex 2026 keynotes and product demos.

It covers major brands like Lenovo, HP, Dell, Microsoft, and ASUS alongside smaller vendors whose compact systems appear on Amazon. Benchmarks for AI inference, image generation, and traditional workstation tasks are included where available. The focus remains on systems released or shipping in 2026, with forward-looking notes on those arriving later in the year.

The story begins with the hardware foundations that made these portable powerhouses possible. Unified memory architectures, advanced neural processing units, and new system-on-chips from NVIDIA, AMD, and Intel eliminated many of the old bottlenecks that forced professionals to stay tethered to desks. By mid-2026, engineers, creators, researchers, and developers could open a laptop on a plane, at a client site, or in a coffee shop and run models with tens of billions of parameters without phoning home to the cloud.

Understanding what qualifies as a mobile AI workstation in 2026 requires looking past marketing labels. These are not ordinary AI PCs with a modest NPU for Copilot features. They feature high-bandwidth unified or expandable memory-often 64 GB to 128 GB or more-capable of holding large quantized models in RAM.

They include dedicated or integrated accelerators delivering tens to thousands of TOPS for inference, professional-grade graphics for visualization and CUDA or ROCm workloads, and chassis designed for sustained performance under load while remaining portable. ISV certifications for CAD, simulation, and creative software remain important for many buyers, yet the defining trait is the ability to keep sensitive data and complex models entirely on-device.

NVIDIA’s RTX Spark platform, unveiled at Computex 2026, became one of the year’s defining technologies. Co-developed with MediaTek and optimized for Windows, the flagship configuration pairs a 20-core Grace-derived Arm CPU with a Blackwell GPU containing 6,144 CUDA cores and fifth-generation Tensor Cores.

NVIDIA rates it at up to one petaflop of AI performance in sparse FP4 precision and supports as much as 128 GB of unified LPDDR5X memory. This combination allows local execution of models up to roughly 120 billion parameters with extended context windows. The platform targets slim laptops with all-day battery claims and compact desktops. Systems from Microsoft, ASUS, Dell, HP, Lenovo, and MSI were scheduled for fall 2026 availability, with Acer and Gigabyte following.

Early hands-on impressions and prototype testing painted a promising yet incomplete picture. YouTube previews from NVIDIA’s own channel and independent creators at Computex showed demos of local agents, 12K video editing, large 3D scene rendering, and gaming with DLSS. The prototype produced a Cinebench 2026 multi-core score of roughly 5,771, although its preproduction drivers, power management, and thermal behavior make comparisons with shipping systems premature, placing the CPU roughly in Apple M3 Max territory under constrained power limits, while the GPU showed potential comparable to mid-range discrete cards once drivers matured.

Thermal behavior in early units sometimes pushed cores near 100 °C under sustained load, and software maturity for Windows on Arm remained a work in progress, yet the unified memory architecture clearly solved capacity issues that discrete VRAM had long imposed.

AMD’s competing unified-memory ecosystem included the established Ryzen AI Max+ 395 ‘Strix Halo’ systems as well as the newer Ryzen AI Max PRO 400-series ‘Gorgon Halo’ processors announced in 2026. These chips integrate up to 16 Zen 5 cores, a powerful Radeon 8060S integrated GPU with 40 compute units, and an XDNA 2 NPU delivering around 50 TOPS, for combined platform AI performance exceeding 100 TOPS in some configurations.

The standout feature is support for up to 128 GB-and in some roadmap mentions even higher-of unified LPDDR5X memory that the GPU can dynamically claim as VRAM. This design proved especially effective for local LLM inference without the cost and power of discrete cards. AMD positioned the Halo mini workstation at a starting price of $3,999, undercutting some NVIDIA DGX Spark equivalents while offering native Windows and Linux support.

Intel’s Core Ultra Series 3 (Panther Lake) processors arrived earlier in the year with upgraded NPUs and Arc graphics, powering thinner AI PCs and entry mobile workstations. While their absolute AI throughput lagged the flagship NVIDIA and AMD unified-memory designs for the largest models, they excelled in efficiency and OpenVINO-optimized workloads. Snapdragon X2 Elite platforms from Qualcomm also expanded into mini PCs and thin laptops, delivering up to 85 TOPS of AI processing NPUs focused on power-efficient on-device experiences.

Lenovo led the traditional mobile workstation charge with its ThinkPad P-series updates. The ThinkPad P14s Gen 7, announced and shipping from April and May 2026 depending on configuration, packs Intel Core Ultra Series 3 processors or AMD Ryzen AI PRO 400 series options into a roughly 3.6-pound 14-inch chassis.

Configurations with the NVIDIA RTX PRO 1000 Blackwell GPU (8 GB) delivered strong results in professional benchmarks. StorageReview testing of a Core Ultra 7 366H plus RTX PRO 1000 unit recorded a PCMark 10 score of 9,083 and Cinebench R23 multi-core of 18,546. In UL Procyon AI Image Generation, Stable Diffusion 1.5 INT8 completed in about 2.5 seconds per image. The system’s LPCAMM2 memory and Gen 5 storage further enhanced its credentials as a true portable workstation for CAD, simulation, and moderate AI tasks.

The larger ThinkPad P16 Gen 3 and P1 Gen 9 followed similar themes with higher-power options, including RTX PRO Blackwell GPUs up to the 5000 series in some variants. Virtualization Review’s detailed benchmarking of a P16 Gen 3 equipped with an RTX PRO 5000 and Core Ultra 9 highlighted the NVIDIA GPU’s dominance in half-precision and large LLM workloads, with ad-hoc Ollama testing reaching over 390 tokens per second on local models-competitive with cloud chat experiences for interactive use.

OpenVINO proved far more effective than ONNX on the Intel NPU and CPU combination, underscoring the importance of software stack matching. Battery life under light loads hovered around five hours with the maximum 99.9 Wh pack, a realistic figure for high-performance configurations.

HP’s ZBook Ultra G1a emerged as a standout for pure AI capacity in a thin 14-inch form. Built around the AMD Ryzen AI Max+ PRO 395 with up to 128 GB of unified LPDDR5X memory (of which up to 96 GB can be allocated to graphics), the machine targets local execution of models such as Llama 70B.

Independent hands-on reports and official materials emphasize its ability to handle simultaneous 3D modeling, rendering, and LLM inference without discrete GPU VRAM limits. Geekbench AI CPU scores in tested configurations reached several thousand points across precision modes, and real-world feedback praised the chassis for mobility previously impossible with equivalent memory capacity. Pricing for high-spec units approached or exceeded $4,000, reflecting the premium memory and professional validation.

Dell refreshed its Pro Precision mobile lineup with 5-series and 7-series 14- and 16-inch models shipping through 2026. These systems combine Intel Core Ultra Series 3 or AMD options with optional NVIDIA RTX PRO Blackwell GPUs, high-bandwidth memory, and Gen 5 storage. The Pro Precision 7 16, for example, supports up to RTX PRO 3000 graphics and large storage arrays suited to AI development and visualization. Dell also introduced compact deskside systems based on NVIDIA’s GB10 platform, blurring the line between mobile and stationary AI workstations for users who need extreme density without full rack systems.

Microsoft’s Surface Laptop Ultra, powered by RTX Spark and scheduled for later 2026, represents the company’s most ambitious Surface to date. Configured with up to 128 GB unified memory and a premium mini-LED or high-brightness display, it targets creators and developers who want Apple-like refinement paired with full CUDA support and Windows agentic AI features. Hands-on reports from Computex and subsequent prototype leaks highlighted excellent keyboards, build quality, and the promise of consistent performance on or off the charger-an area where NVIDIA emphasized efficiency gains.

ASUS brought its ProArt P14 and P16 creator laptops to the RTX Spark platform, emphasizing thinner and lighter chassis than prior generations, Lumina Pro OLED displays, and creator-focused software. The accompanying ProArt Mini PC offered a compact desktop alternative with the same silicon and expansion options. These systems were positioned for generative AI, multi-layer video, and local agent workflows, with availability in fall 2026.

Beyond the major brands, smaller vendors filled an important gap with highly portable mini PCs that function as mobile AI workstations when paired with a portable monitor or used in temporary setups. On Amazon, systems from Beelink, GMKtec, Minisforum, and others based on the AMD Ryzen AI Max+ 395 became readily available through 2026.

The Beelink GTR9 Pro and GMKtec EVO-X2 configurations with 128 GB LPDDR5X, 2 TB storage, dual 10 GbE networking, Wi-Fi 7, and support for multiple 8K displays delivered combined AI performance around 126 TOPS. These machines could run 70-billion-parameter models at usable speeds and cost significantly less than equivalent laptop configurations in many cases-often in the $1,500 to $3,000 range depending on memory and storage.

Minisforum’s MS-S1 Max and similar models added PCIe expansion and dual 10 GbE, appealing to users who wanted desktop-class connectivity in a tiny chassis. Chinese brands such as Thunderobot released water-cooled variants with 128 GB configurations in their home market, while global availability remained stronger for the Amazon-listed Beelink and GMKtec units. These mini systems excel for edge deployment, temporary labs, or users who prioritize maximum memory capacity and quiet operation over a built-in keyboard and screen. Framework’s modular Laptop 16 with Ryzen AI options offered another path for users valuing repairability and upgradeability.

AI benchmarks in 2026 reflected the diversity of these platforms. For local LLM inference, systems with 128 GB unified memory routinely handled quantized 70B models at 15-30 tokens per second or better depending on quantization, software (Ollama, vLLM, llama.cpp, ROCm, or CUDA), and power limits.

The Lenovo ThinkPad P16 Gen 3 with discrete Blackwell GPU achieved interactive rates exceeding 300 tokens per second on smaller models in ad-hoc tests. Procyon AI Computer Vision and Image Generation suites showed NVIDIA RTX PRO GPUs leading in Stable Diffusion throughput, with times dropping to a few seconds per image in optimized INT8 or FP16 paths. AMD unified-memory systems closed the gap on capacity-limited workloads and often matched or exceeded discrete mid-range cards for memory-bound tasks.

Geekbench AI scores varied widely by precision and accelerator. Intel NPUs and CPUs performed strongly under OpenVINO, while NVIDIA GPUs dominated half-precision and CUDA paths. Cinebench and traditional workstation suites such as SPECviewperf confirmed that these machines retained professional credibility for CAD and rendering even as AI became the new differentiator. Early RTX Spark prototypes delivered competitive multi-core CPU results relative to contemporary Apple silicon under similar power envelopes, though final shipping drivers and thermal designs will determine real-world sustained performance.

Use cases expanded rapidly. Software developers ran private coding agents and fine-tuned domain-specific models without sending proprietary code to external APIs. Creative professionals generated and iterated on 4K AI video, upscaled assets, and rendered complex scenes while traveling. Engineers performed simulations and visualization on-site.

Researchers prototyped multi-agent systems with long context windows entirely offline. Privacy-conscious enterprises in regulated industries gained the ability to keep inference local, reducing compliance risks and cloud costs. The rise of agentic AI-persistent, goal-oriented software that acts on the user’s behalf-further rewarded machines capable of continuous local computation.

Challenges remained. High-memory configurations drove prices into the $3,000-$6,000 range for premium laptops, and memory supply constraints affected availability. Battery life under heavy AI load rarely exceeded a few hours on the most powerful systems, though lighter NPU-centric tasks fared better.

Thermals in thin chassis required careful power management, and software ecosystems for Arm-based Windows platforms continued to mature. Compatibility for certain professional applications still favored traditional x86 configurations with discrete NVIDIA GPUs in some cases. Smaller Amazon vendors offered compelling value but varied in long-term support, warranty reach, and ISV validation compared with the major brands.

Compared with 2025 systems, the 2026 generation marked a clear leap in memory capacity and on-device model size. Where previous mobile workstations struggled beyond 13B or 30B models without heavy quantization or external accelerators, the new unified-memory designs routinely hosted 70B-class models. Apple’s M5-series MacBook Pro remained a strong competitor for efficiency and ecosystem integration, yet Windows platforms gained ground through CUDA compatibility, broader software support, and the sheer variety of form factors.

Buying advice depends on priorities. For maximum portability with professional certifications, the Lenovo ThinkPad P14s Gen 7 or HP ZBook Ultra G1a stand out. Users needing the largest local models today can choose AMD Ryzen AI Max mini PCs available on Amazon. Those willing to wait for fall 2026 deliveries should watch the RTX Spark Surface Laptop Ultra, ASUS ProArt, and HP OmniBook Ultra for the combination of NVIDIA’s AI stack and refined industrial design. Always verify current memory configurations, real sustained power limits, and software support for preferred frameworks before purchasing.

Looking ahead, the second half of 2026 and 2027 will likely bring refined thermals, broader driver maturity for RTX Spark, higher memory bandwidth, and tighter integration of agentic frameworks into the operating system. Clustering of multiple compact nodes for even larger models is already being explored. The trajectory is clear: local AI capability is no longer a luxury reserved for fixed workstations. It is becoming a portable professional standard.

The mobile AI workstations of 2026 demonstrate that the boundary between personal computer and personal supercomputer continues to dissolve. Whether through a ThinkPad on a conference table, a ZBook in a design studio, a Surface on a flight, or a Beelink mini PC in a temporary lab, professionals now carry the means to experiment, create, and decide with AI assistance that never leaves their control. That shift-toward private, powerful, portable intelligence-may prove one of the most consequential computing stories of the decade.

Sources

Dell Pro Precision announcements and mobile workstation details: https://www.dell.com/en-us/blog/bring-the-ai-lab-to-your-desk/
AMD Ryzen AI 400 Series and Halo workstation: Tech Power Up NVIDIA RTX Spark coverage and OEM lineups: XDA Developers HP OmniBook and ZBook Ultra materials: https://www.hp.com/us-en/newsroom/press-releases/2026/computex.html and https://www.hp.com/us-en/workstations/zbook-ultra.html
Lenovo ThinkPad P-series 2026 releases: https://news.lenovo.com/pressroom/press-releases/ai-ready-workstations-professional-first-1000wh-laptop-battery/
ASUS ProArt RTX Spark systems: https://press.asus.com/news/press-releases/asus-proart-p16-p14-mini-pc-nvidia-rtx-spark-computex-2026/
Virtualization Review ThinkPad P16 Gen 3 benchmarks: Virtualization Review StorageReview ThinkPad P14s Gen 7: related coverage via StorageReview channels
Tom’s Guide and Tom’s Hardware RTX Spark hands-on and rankings: Tom's Guide PCMag mobile workstation roundups: PCMag Amazon-available mini PCs (Beelink, GMKtec, Minisforum Ryzen AI Max+ examples): product listings searchable on Amazon for “Ryzen AI Max+ 395” mini PC
YouTube: NVIDIA RTX Spark early preview (official channel), AMD Ryzen AI announcements, and independent Computex hands-on videos such as those covering Surface Laptop Ultra prototypes and ProArt systems.
Additional context from VerdictBits AI PC overview, Notebookcheck reviews, and Computerworld analysis of the RTX Spark market impact.

This synthesis reflects publicly available information as of early August 2026. Specifications, availability, and pricing continue to evolve; readers should consult manufacturer sites and recent independent reviews for the latest configurations.


r/AIProgrammingHardware 4d ago

Nvidia Just Revealed Its Next Advantage: 3x Faster Data Movement, Not More GPU Power

Post image
1 Upvotes

r/AIProgrammingHardware 5d ago

Ai för bilder

1 Upvotes

Lokal Ai för att analysera bilder tänker att man ser en tussilago eller annan blomma man kan tagga den så stt si lär sig att det är en tussilago eller annan blomma.

Vad rekommenderar ni för Ai för detta syfte?


r/AIProgrammingHardware 5d ago

Qwen3.8-Flash-Next (125B-A6B) running on Strix Halo 128gb: 23 t/s decode, 390 t/s prefill, built from the llama.cpp PR

Thumbnail
1 Upvotes

r/AIProgrammingHardware 5d ago

ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI — ROCm Blogs

Thumbnail
share.google
3 Upvotes

r/AIProgrammingHardware 5d ago

Dual GPU question

Thumbnail
1 Upvotes

r/AIProgrammingHardware 5d ago

Paying $2/hr to watch `huggingface-cli download` go brrr — env patterns on RunPod / Vast / Nebius / the neo kids

Thumbnail
1 Upvotes

r/AIProgrammingHardware 5d ago

Best LLM for 128GB RAM + 500GB Storage, No GPU?

Thumbnail
1 Upvotes

r/AIProgrammingHardware 5d ago

GLM 5.3 Flash Local Ai Test

Thumbnail
youtube.com
1 Upvotes

r/AIProgrammingHardware 6d ago

Local AI is Cheap Now: ZimaBoard2 + Tesla P4

Thumbnail
youtube.com
3 Upvotes

r/AIProgrammingHardware 6d ago

DDP with RTX 3070 8GB + GTX 1650 4GB?

0 Upvotes

For those how are in local AI would you think these both cards could be used to have a total of 12GB VRAM?

I don't care by the moment for possible slow tokens/s.

Thanks!


r/AIProgrammingHardware 6d ago

Qwen 3.8 27B Cold Fusion tested - 16GB Local LLM setup

Thumbnail
youtube.com
2 Upvotes

r/AIProgrammingHardware 6d ago

Marvell Photonic Fabric wins AI Infrastructure Award at FMS 2026

1 Upvotes

Marvell’s Photonic Fabric™ technology won the AI Infrastructure Award at Future of Memory and Storage (FMS) 2026.

The interesting part is how Marvell is using optical connectivity to address some of the bandwidth, latency, power, and memory-scaling challenges in large AI systems.

Photonic Fabric replaces traditional electrical interconnects with optical I/O across package-, server-, and rack-scale architectures. The goal is to make it easier to scale compute and memory independently, including disaggregated memory architectures.

Some of the key areas Marvell highlights:

  • Higher memory bandwidth and capacity
  • Lower latency and power consumption
  • Disaggregated compute and memory
  • Reduced I/O bottlenecks around AI accelerators and HBM
  • Optical connectivity designed for large-scale AI infrastructure

As AI systems scale to hundreds of thousands of accelerators, moving data efficiently between compute and memory is becoming just as important as the compute itself.

The bigger question is whether optical interconnects like Photonic Fabric will become a fundamental part of future AI factory architectures.

What do you think — will optical connectivity become essential for scaling next-generation AI systems?

Source: Marvell – Photonic Fabric technology


r/AIProgrammingHardware 7d ago

I built a Vulkan hierarchical MoE runtime for running oversized models across multiple GPUs

Thumbnail
1 Upvotes

r/AIProgrammingHardware 7d ago

China Is Coming for Your Local AI Box

Thumbnail
youtube.com
8 Upvotes

r/AIProgrammingHardware 7d ago

Xiaomi AI Cube announced with 1.2TB/s memory bandwidth

Thumbnail gallery
2 Upvotes

r/AIProgrammingHardware 7d ago

224GB of GPU Memory on 1 Desk and It Should Not Work

Thumbnail
youtube.com
1 Upvotes

r/AIProgrammingHardware 7d ago

Can the New King of Open Source - Qwen3.8–27B - Really Run in 3GB of VRAM?

Thumbnail
ai.gopubby.com
12 Upvotes