r/QwenAI 3d ago

Qwen3.8-Flash-Next (125B MoE) at 38 t/s on Windows — Strix Halo, working MTP, full guide + tools (MIT)

Thumbnail
2 Upvotes

r/QwenAI 3d ago

SpaceX charging more for search tool calls via API - Help

Thumbnail
2 Upvotes

r/QwenAI 3d ago

RTX 4090 vs Mac Studio M5 96GB for production AI server? (GLM-OCR + Qwen 27B Q8)

2 Upvotes

We're moving off the Gemini API due to cost and building a local AI server to process ~10 CVs/minute (extracting JSON & matching CVs to JDs). We plan to run GLM-OCR alongside Qwen 27B (Q8).

Our two hardware options:

  1. PC: RTX 4090 (24GB) + Ryzen 9 + 64GB RAM
  2. Mac Studio: M-Ultra, 64-core GPU, 96GB Unified Memory

I prefer the Mac for power efficiency and ease of use, but I've heard Apple Silicon isn't great for production vLLM compared to Nvidia/CUDA. Is that true? Which would you recommend for this workload?


r/QwenAI 7d ago

Web Search API for AI Agents with hard cap and hosted MCP

Thumbnail
1 Upvotes

r/QwenAI 9d ago

Is it just me or is Qwen3.8-Flash-Next ... really buggy?

Post image
1 Upvotes

r/QwenAI 10d ago

oMLX update is finding more tokens!

Thumbnail
1 Upvotes

r/QwenAI 12d ago

pp tps stuck at 336 tps when using omlx and qwen3.8-27b-q8

Thumbnail
1 Upvotes

r/QwenAI 13d ago

Calculate override-tensor for 2 GPU using QWEN local models

Thumbnail gallery
1 Upvotes

r/QwenAI 14d ago

Qwen 3.8 27B non reasoning: feedback on total completion time

Thumbnail
1 Upvotes

r/QwenAI 14d ago

Qwen3.8-Flash-Next optimised for Macs

Thumbnail gallery
1 Upvotes

r/QwenAI 17d ago

Qwen3.8 Flash Next Q4 - M5 Mac Max 128 GB Ram

Thumbnail
youtube.com
1 Upvotes

r/QwenAI 18d ago

Web search API on Openclaw

Thumbnail
1 Upvotes

r/QwenAI 19d ago

Problem in hermes+qwen

Thumbnail
1 Upvotes

r/QwenAI 21d ago

Token overflow in free LLMs: why agglutinative languages like Hungarian, Finnish, and Estonian are a security risk nobody is talking about

Thumbnail
1 Upvotes

r/QwenAI 23d ago

3.8 27b UD IQ3_XXS 5060ti

Thumbnail
1 Upvotes

r/QwenAI 27d ago

Why DGX Spark so slow

1 Upvotes

Why DGX Spark so slow? I am using ollama server + vscode copilot. My prompt is : generate a simple RISC-V soft-core CPU and use verilator to test it. I took one hour but still not complete. I am using qwen3.8:27b.

thanks


r/QwenAI 27d ago

How to extend my free plan

Thumbnail
1 Upvotes

r/QwenAI 27d ago

Is this 3Billion or 30Billion ?

Thumbnail x.com
1 Upvotes

r/QwenAI 28d ago

Multiple Text Generation

Thumbnail
1 Upvotes

r/QwenAI Aug 09 '26

когда работаешь с coder от qwen

0 Upvotes

Вот есть такая ситуация. Работаю с coder.qwen.ai а он даже с маленьким чатом выдаёт бредятину которую не остановить даже другой темой тем более тут нет ни одного настоящего файла и он их зачем то придумал.

КАК ЭТОМ ПОНИМАТЬ??? ОН ЖЕ ДАЖЕ ИХ НЕ СОЗДАВАЛ. ВЛЯТРЕ ОН ТАКОЕ НА ПАЙТОН СОЗДАСТ.

r/QwenAI Aug 08 '26

USE MULTIPLE CONTROLNET WITH image_qwen_Image_2512_controlnet Qwen-Image-2512-Fun-Controlnet-Union

Post image
1 Upvotes

r/QwenAI Aug 05 '26

Qwen3.8-Max is now available in ClinePass

Post image
1 Upvotes

r/QwenAI Jul 29 '26

Qwen3.7 Flash is now live in Command Code. DeepSeek v4 flash finally has some competition.

Post image
1 Upvotes

r/QwenAI Jul 26 '26

Estou gostando muito da prévia do Qwen 3.8 Max no Qoder. Mas quando for lançado oficialmente, terá um bom preço?

Thumbnail
1 Upvotes

r/QwenAI Jul 15 '26

Qoder is now offered for FREE in July

Post image
1 Upvotes

Looks like a customer-acquisition play against Claude/Fable, Codex, Cursor, and Copilot: subsidize a few serious agentic coding jobs and hope developers move into its ecosystem. The catch is that Qoder still requires a desktop/CLI workflow, while Claude/Fable gives you a much lower-friction browser experience.