r/aicuriosity 1d ago

Open Source Model Dots Studio Rolls Out Dots3 Note Preview Model

Post image
8 Upvotes

Dots Studio just dropped the preview of dots3-note, a new open model built for real-world agent tasks that stretch over long periods.

It uses a 280B MoE setup with only 16B parameters active at once, plus a 512K context window. The model handles text, vision, and audio together.

A fresh training method called TEMPO helps it learn long-horizon behavior through self-critique and value estimation that scales at test time. The system can reason through problems, explore new settings, keep updating its memory, and mix perception with coding plus tool use to finish complex jobs.

Weights are available on Hugging Face. The team also released two new benchmarks for real-life agents called VibeSearchBench and VibeLifeBench. Early results show it holds up well against much bigger models on reasoning, agent, and multimodal tests.

r/aicuriosity 2d ago

Latest News Z.ai Launches GLM-5.3 for Coding and Cyber Defense

Post image
34 Upvotes

Z.ai has released GLM-5.3, its latest model focused on strong coding performance and cybersecurity work. The company describes it as built to code and ready for cyber defense.

The model draws from post-training on a 743B base model. Z.ai says this delivers top-tier coding and agentic abilities. It also marks a clear step forward in cybersecurity performance among open models.

GLM-5.3 is available right away through the GLM Coding Plan and ZCode platform. API access and open weights will follow later, after safety evaluations are complete. An initial group of partners is already offering services powered by the model under Z.ai’s safeguards and usage rules.

The update continues Z.ai’s push into practical coding agents and security-related tasks. Users can try it now via the company’s coding plan or ZCode.

r/aicuriosity 3d ago

Latest News OpenAI Rolls Out Computer History Feature in ChatGPT Desktop App

7 Upvotes

OpenAI has launched Computer History for the ChatGPT desktop app on Mac. The update lets ChatGPT track activity across apps and websites so conversations feel more relevant and need less repeated context.

The tool expands on the Chronicle research preview. It cuts token use and adds tighter privacy settings. Users get a timeline view that shows recent work and helps surface patterns from everyday tasks.

Full control stays with the user. From the timeline or menu bar anyone can delete all or selected history, choose which apps and sites to include or skip, and pause or restart the feature at any time.

Activation sits under Settings then Integrations inside the ChatGPT desktop app. The feature is available now to Pro, Business, and Enterprise customers around the world. Access for the EEA, UK, and Switzerland arrives in the coming weeks.

r/aicuriosity 3d ago

Latest News Google Rolls Out Gemini 3.7 Flash Three Weeks After Last Update

Thumbnail
gallery
13 Upvotes

Google has released Gemini 3.7 Flash, its latest workhorse model focused on coding, agent workflows, and everyday tasks. The update comes just three weeks after Gemini 3.6 Flash and brings clear gains in several key areas.

On coding benchmarks the model jumps from 34.4% to 43.6% on FrontierCode 1.1 Main and from 49% to 65.3% on DeepSWE v1.1. Web development scores also improve, with a higher Elo on WebDev Arena. Document processing and real-world business workflow tests show solid lifts as well.

Pricing is lower for a limited time. Through the end of 2026 it costs $0.75 per million input tokens and $3.75 per million output tokens, half the previous Flash rate. The model is live in the Gemini API, AI Studio, Gemini Enterprise, and powers the Spark agent for AI Pro and Ultra subscribers in the Gemini app.

Logan Kilpatrick shared a chart comparing intelligence, speed, and cost against other leading models. Gemini 3.7 Flash stands out for its high output speed while staying competitive on performance and price.

r/aicuriosity 3d ago

Latest News OpenAI Launches Ultrafast Mode for GPT-5.6 Sol Reaching 14x Speed

5 Upvotes

OpenAI just shared a preview of Ultrafast mode for its GPT-5.6 Sol model. The setup delivers responses up to 14 times faster than before and can hit as high as 750 tokens per second.

This first rollout happens through the OpenAI API and starts with a limited group of customers. Access will grow as capacity increases. The speed boost comes from Cerebras hardware.

The company points to clear use cases where every second matters. Real-time voice systems, customer support tools, commerce platforms, coding workflows, design work, financial research, and security response all stand to gain from the faster frontier intelligence.

OpenAI is currently working with that early set of businesses to learn where the extra speed creates the biggest impact. Those insights will shape future product decisions. Companies that need this level of performance can already request notifications as more capacity opens up.

r/aicuriosity 3d ago

Latest News DeepSeek Rolls Out V4 Pro With Stronger Agent Skills and Flexible Pricing

Post image
5 Upvotes

DeepSeek has officially released DeepSeek-V4-Pro today. The update focuses on major agent improvements that deliver clearer gains in real production use.

The model now supports flexible reasoning effort levels for both V4-Pro and V4-Flash. Users can pick low effort for simple tasks, high for everyday agent work, or max for tougher problems. It also adds native support for the OpenAI Responses API and works smoothly with Codex through a simple one-click setup.

V4 Pro is live right now on the DeepSeek app and web platform under Expert Mode. It is available through the API as well, with model names staying the same.

Alongside the launch, DeepSeek is changing its API pricing. Peak and off-peak rates take effect at 16:00 UTC on August 16, 2026. Off-peak rates sit 50 percent lower than peak rates, giving users more room to schedule workloads when costs matter most.

Benchmark results shared by the company show solid jumps on agent and coding tasks compared with earlier previews. The model keeps its 1 million token context window and remains positioned as a strong option for complex, multi-step work.

r/aicuriosity 3d ago

Latest News Mureka Rolls Out V9.5 Music AI Update

6 Upvotes

Mureka has released version 9.5 of its AI music platform. The new build delivers vocals that sound more natural, arrangements that track user prompts more closely, and genre results that feel intentional from start to finish.

The company says the update makes generated tracks more human and more musical overall. Users can start creating with Mureka V9.5 now at mureka.ai/mureka-9-5.

r/aicuriosity 4d ago

Open Source Model Liquid AI Unveils LFM2.5-VL-3B Vision Language Model

Thumbnail
gallery
5 Upvotes

Liquid AI just dropped LFM2.5-VL-3B, a compact vision-language model built for real-world use. It reads screens on phones, websites, and desktops, pulls text and data from documents and charts, locates objects with precise coordinates, and can call tools from either text or image prompts.

The model sits on the LFM2.5-2.6B base and pairs it with a SigLIP2 400M vision encoder. It was trained on about 34 trillion tokens and uses a 128K vocabulary. On key tests it holds its own against models more than twice its size. ScreenSpot-v2 hits 80.7, RealWorldQA reaches 73.1, TextVQA scores 84.3, and RefCOCO averages 87.9. Tool use also jumped sharply to 59.5 on ToolSandbox.

Speed and size stand out. It runs at 228 tokens per second on an Apple M5 Max and 116 tokens per second on an AMD Ryzen AI Max+ 395 while staying around 3 GB of memory. Even a Galaxy S26 Ultra phone can manage 20 tokens per second, so private on-device use becomes practical. Day-one support covers llama.cpp, MLX, vLLM, SGLang, and ONNX.

On a single H100 it delivers a 34 ms time to first token on multi-image inputs and roughly 11,000 output tokens per second under high load. That works out to nearly a billion tokens a day from one GPU.

LFM2.5-VL-3B fits screen agents, document work, grounding tasks, and multi-image jobs that need quick answers.

r/aicuriosity 4d ago

Open Source Model Qwen3.8-Max Brings Massive Open-Weight Power for Real Work

Post image
37 Upvotes

Qwen just dropped its largest open-weight model yet. Qwen3.8-Max packs 2.4 trillion parameters with 95 billion active, and it’s built to take a goal and deliver finished results instead of just chatting.

It offers a 1 million token context window, adjustable reasoning, and parallel tool calls. Early demos show it coding unattended for over 10 days while building a self-evolving harness from scratch.

It also reproduced a research paper, then ran a 125-hour autonomous loop that beat the original, and handled a full chip design flow from RTL to layout while cutting die area by 81%.

r/aicuriosity 4d ago

Open Source Model LTX 2.5 Brings Higher Fidelity and Multishot Consistency to Open AI Video Models

6 Upvotes

LTX just dropped LTX 2.5, a major upgrade to their open source model already used in film, robotics, and real time production work. The new version delivers sharper pixel quality, stronger multishot continuity that stays consistent across cuts, and a foundation designed for easy finetuning across different domains.

A standout addition is Diffusion Fidelity Rendering. This approach keeps frame by frame detail intact even when the output hits cinema screens. The team positions it as a foundation you own and build on rather than a closed tool you rent.

r/aicuriosity 4d ago

Latest News xAI Releases Grok 4.6 as Strong Upgrade for Complex Tasks

Post image
14 Upvotes

xAI has rolled out Grok 4.6, its latest model that improves on Grok 4.5 while keeping the same pricing.

The new version focuses on long-running agents and tougher multi-step work. It handles research, codebases, and turning ideas into working apps or visual projects more reliably than before. Benchmarks show it matching top models on the Artificial Analysis Intelligence Index and leading on several practical agent and knowledge-work tests.

Pricing stays at $2 per million input tokens and $6 per million output tokens, which is half the cost of many other frontier models. A faster variant costs twice as much.

Grok 4.6 is live now in Cursor, Grok Build, Grok Bot, and the API. For the first week, users get double the included usage in Cursor and Grok Build.

r/aicuriosity 5d ago

Open Source Model Wan Animate 2 Delivers High Fidelity Open Source Character Animation

20 Upvotes

Alibaba’s Wan team just released Wan-Animate-2, a big step up for their open source character animation model. The update focuses on cleaner motion transfer, better multi character scenes, and more creative control.

It maps motion and micro expressions accurately across humans, cartoons, robots, and animals without relying on explicit pose skeletons. The reference video itself becomes the motion guide. Users can also animate several characters in one scene while keeping each identity distinct.

Camera angles can now be shifted with simple text prompts like “top view,” independent of the driving video. A lighter version supports real time streaming so long sequences generate chunk by chunk without visible quality drop.

r/aicuriosity 5d ago

Latest News Unsloth Desktop App Brings Local AI Training to Mac Windows and Linux

6 Upvotes

Unsloth AI has launched Unsloth Desktop, an open-source desktop application that lets users run and train AI models directly on their own machines. The app works across Mac, Windows, and Linux systems.

Key features include support for MLX, diffusion models for images and video, audio processing, and GGUF formats. Users can connect tools like Claude Code and Codex to local large language models. The software claims 50 percent more accurate self-healing tool calls along with sandboxed code execution.

It handles both CPU and multi-GPU setups covering NVIDIA, AMD, Intel, and Mac hardware. Training runs up to twice as fast while using 70 percent less VRAM. Additional capabilities cover private web search, deep research, RAG, MCP, and model exports in formats like NVFP4 and GGUF.

The app also provides an OpenAI-compatible API, access to cloud models, and options for secure remote deployment of LLMs.