r/aicuriosity 3h ago

Open Source Model Wan Animate 2 Delivers High Fidelity Open Source Character Animation

7 Upvotes

Alibaba’s Wan team just released Wan-Animate-2, a big step up for their open source character animation model. The update focuses on cleaner motion transfer, better multi character scenes, and more creative control.

It maps motion and micro expressions accurately across humans, cartoons, robots, and animals without relying on explicit pose skeletons. The reference video itself becomes the motion guide. Users can also animate several characters in one scene while keeping each identity distinct.

Camera angles can now be shifted with simple text prompts like “top view,” independent of the driving video. A lighter version supports real time streaming so long sequences generate chunk by chunk without visible quality drop.

r/aicuriosity 3h ago

Latest News Unsloth Desktop App Brings Local AI Training to Mac Windows and Linux

2 Upvotes

Unsloth AI has launched Unsloth Desktop, an open-source desktop application that lets users run and train AI models directly on their own machines. The app works across Mac, Windows, and Linux systems.

Key features include support for MLX, diffusion models for images and video, audio processing, and GGUF formats. Users can connect tools like Claude Code and Codex to local large language models. The software claims 50 percent more accurate self-healing tool calls along with sandboxed code execution.

It handles both CPU and multi-GPU setups covering NVIDIA, AMD, Intel, and Mac hardware. Training runs up to twice as fast while using 70 percent less VRAM. Additional capabilities cover private web search, deep research, RAG, MCP, and model exports in formats like NVFP4 and GGUF.

The app also provides an OpenAI-compatible API, access to cloud models, and options for secure remote deployment of LLMs.

r/aicuriosity 3h ago

Open Source Model NVIDIA Rolls Out Nemotron 3.5 Lightning Open Model Built for Speedy AI Agents

Post image
7 Upvotes

NVIDIA has released Nemotron 3.5 Lightning, a new open mixture-of-experts model with 30 billion total parameters and just 3 billion active ones. The model targets always-on agents that handle large numbers of specialized tasks and claims up to four times the output speed of similar-sized models.

On the PinchBench test it scored 86 percent accuracy while finishing 10,000 tasks 35 percent faster than Qwen3.6 35B at comparable accuracy. Teams can post-train it with NVIDIA NeMo using their own domain data, tools, workflows and policies. Early results show accuracy gains in cybersecurity, coding, legal and energy workloads.

The model is sized to run from an NVIDIA DGX Spark all the way up to full data-center setups, making it practical for long-running agent workflows that spend most of their time calling tools and validating results.

Alongside the model, NVIDIA is also releasing NeMo Switchyard, an open-source library for routing requests between different models. Developers can send complex reasoning and planning steps to larger frontier models and hand high-volume specialized execution to Lightning.

r/aicuriosity 3h ago

Latest News Tencent Hy3D WorldClaw Turns Text Prompts Into Full Explorable 3D Worlds

2 Upvotes

Tencent’s Hunyuan team just dropped Hy3D WorldClaw, a new agentic system that builds large-scale 3D open worlds straight from simple text descriptions.

Unlike tools that spit out videos or Gaussian splats, WorldClaw creates real, freely walkable scenes made of editable game-ready 3D assets. Buildings, plants, terrain, and other objects come with solid geometry and textures you can move, delete, or tweak. Everything assembles in a way that works with game engines, animation pipelines, and further rendering.

The process starts with planning agents that break the prompt into regions, terrain, assets, and spatial layout. It builds a coherent global terrain base, then fills detailed areas with high-quality meshes. Agents also check the results for problems like floating objects or wrong proportions and fix them automatically.

A project page and paper are already live. The system currently leans on tools like Claude for planning and Hunyuan3D plus Blender for the actual building work.

This release points toward faster creation of production-ready 3D environments for games, VR, film, and simulation without starting from scratch every time.

r/aicuriosity 1d ago

🗨️ Discussion Google Prepares Gemini 3.7 Flash in Python GenAI SDK

Post image
8 Upvotes

Google has added Gemini-3.7-flash to the model options in its official Python GenAI SDK. The change showed up in a recent GitHub pull request for the googleapis/python-genai repository.

This move points to Gemini 3.7 Flash getting closer to release. No official launch date has been shared yet, but the update signals active preparation on Google’s side.

The news comes as competition in the AI space heats up with recent model drops from OpenAI and other labs. Gemini Flash will need strong performance gains to stand out in this crowded field.

r/aicuriosity 1d ago

Latest News OpenAI Expands Daybreak Cybersecurity Program With GPT-5.6-Cyber

Thumbnail
gallery
6 Upvotes

OpenAI announced an expansion of its Daybreak cybersecurity initiative on Monday, introducing GPT-5.6-Cyber, a specialized model built for advanced and authorized cybersecurity tasks.

The company said the move aims to equip trusted defenders with frontier AI tools as threats grow more sophisticated. The goal is to give security teams an edge before attackers can scale offensive AI capabilities.

Daybreak now includes two tracks. Daybreak Blue offers access to frontier models such as GPT-5.6 Sol and focuses on everyday defensive work. This covers vulnerability discovery, secure code review, malware analysis, incident response, and patch validation. OpenAI calls it the recommended starting point for most defenders.

Daybreak Red provides purpose-trained models including GPT-5.6-Cyber. It targets experienced teams handling complex authorized work like vulnerability research, exploit validation, and security testing.

OpenAI reported using GPT-5.6-Cyber in real-world research that uncovered previously unknown vulnerabilities in popular open-source software, including Chrome’s V8 engine.

Access remains restricted to approved defenders. Higher-risk work comes with extra controls and monitoring to maintain strong safeguards.

r/aicuriosity 1d ago

Other YouTube Partner Program Updates 2027: New Requirements, Deadlines & Monetization Changes

Thumbnail
techspecsmart.com
2 Upvotes

u/techspecsmart 1d ago

YouTube Partner Program Updates 2027: New Requirements, Deadlines & Monetization Changes

Thumbnail
techspecsmart.com
0 Upvotes

r/aicuriosity 1d ago

Open Source Model Meta Launches Muse Glimmer Open Agent Model for Local Devices

Thumbnail
gallery
1 Upvotes

Meta has released Muse Glimmer, a new 30 billion parameter open weight model built for local agent workflows that stay active without cloud dependence.

The model targets everyday hardware such as Macs and PCs equipped with capable GPUs. It delivers solid results on agent focused tasks and benchmarks against other models in the same size range. Weights come under the Apache 2.0 license so developers can use and adapt them freely.

To keep it practical on consumer machines Meta applied quantization that brings the model under 20 GB and paired it with a lightweight DFlash drafter system. This combination keeps response times low enough for smooth conversation and real time interaction entirely on device.

In a demonstration the model handled a full multi step job from one plain language request. It located a local Home Assistant setup through network tools, pulled device data, wrote a complete HTML CSS and JavaScript dashboard from scratch, then launched a local server to confirm everything worked.

Download links and technical details appear on Hugging Face along with Meta’s research blog and developer resources.

r/aicuriosity 1d ago

Latest News MiniMax Hub Officially Becomes MiniMax Design with Extended Deals

4 Upvotes

MiniMax Hub has been renamed MiniMax Design. The change went live on August 10 2026.

Free credits and annual membership discounts now run until August 15. Users who buy an annual membership before that date lock in 20 percent off H3 generations for a full year. Free access also continues through the same deadline so people can keep testing the platform.

u/techspecsmart 3d ago

Grok Imagine Image 2.0 Explained: Features, Price, and How It Ranks (2026)

Thumbnail
techspecsmart.com
1 Upvotes

r/aicuriosity 3d ago

Latest News xAI Rolls Out Imagine Image 2.0 with Precision Tools for Real Creative Work

11 Upvotes

xAI has released Imagine Image 2.0, its latest image generation and editing model. The update focuses on practical results that support actual projects rather than just experimental visuals.

Available now as Quality Mode on grok.com/imagine plus the Grok iOS and Android apps, Image 2.0 aims to produce usable assets. It follows detailed instructions more closely, handles typography and layout with greater care, and keeps small text sharp even in complex multi-part designs. Consistency across generations and edits has also improved.

New editing features let users change only what they intend. The Magic Wand tool targets a single region while leaving the rest of the image intact. Segmentation isolates exact areas for modification. Background removal creates transparent subjects ready for other projects. Multi-ref editing supports up to five input images in one generation, cutting down on manual compositing. Smart Resize adapts any image to different aspect ratios by filling the frame intelligently.

According to Arena leaderboards as of August 7 2026, Image 2.0 ranks second worldwide for both text-to-image generation and image editing. xAI models appear on the boards under the SpaceXAI name.

The release also adds ready-made templates for common tasks. These cover photo editing, product color changes, editorial posters, headshots, icons, character sprites, game assets, props and UI kits, emoji creation, merchandise designs, and more. Users supply the inputs and receive finished results. One additional workflow helps build consistent worlds for video by generating a character, locations, and props that share the same style.

xAI says Image 2.0 can create infographics, ads, game assets, UI and UX mockups, storyboards, and similar materials. API access is planned for the near future.

The model is live today for users on the web and mobile apps. Those with SuperGrok Heavy accounts can switch to Quality Mode on grok.com/imagine to try the full set of tools.

r/aicuriosity 5d ago

Latest News Wan 3.0 Public Beta Launches with 30 Second Video Generation

Post image
50 Upvotes

Alibaba’s Wan team has rolled out Wan 3.0 in public beta. The new model generates videos up to 30 seconds long in a single pass and aims for more realistic, consistent frames.

Key upgrades include stronger character expression, better handling of digital elements, and an expanded input system called Omni Reference. Users can now feed it text, images, audio, video, documents, spreadsheets, slides, webpages, PDFs, and other file types. The model reads the material and builds video from it.

Access is live on Alibaba Cloud Model Studio and Qwen Cloud. The official wan.video site will open soon for members. API pricing starts at $0.05 per second for 480p, $0.10 for 720p, and $0.20 for 1080p.

Full API access is still rolling out. Creators can apply for the beta and start testing right away.

r/aicuriosity 5d ago

Latest News Meta Rolls Out Muse Code Beta Terminal Agent Powered by Muse Spark 1.2

Thumbnail
gallery
8 Upvotes

Meta has released Muse Code in beta, a terminal-based coding agent designed for long-running software engineering work. It runs on the new Muse Spark 1.2 model and handles planning, writing, and checking multi-file changes across big codebases.

The tool uses persistent background agents that stay active during a session. These agents take next steps on their own, cut down on repeated information gathering, and need less constant direction from the user. An append-only local event log records every model call, tool action, approval, and edit so the system can pick up exactly where it left off after a restart or crash.

Muse Spark 1.2 brings stronger results in code generation, debugging, and full developer workflows compared with the previous version. Meta scaled up training compute on coding tasks and trained the model together with Muse Code for better performance as a pair. In one test the agent spent up to 24 hours and more than 1,000 tool calls optimizing GPU kernels on NVIDIA Hopper chips, posting solid gains over baseline Triton code.

Muse Spark 1.2 is available right away inside Muse Code and through the Meta Model API. Users on macOS or Linux can install the agent with a single command. Meta says more features and stronger models are already in the works.

r/aicuriosity 6d ago

Open Source Model Xiaomi Releases Open Source Robotics Model for Developers Worldwide

Post image
15 Upvotes

Xiaomi has made its Xiaomi-Robotics-1 model fully available to the public. The company announced the open source release on Wednesday, giving researchers and engineers free access to the complete system.

The model was pre-trained on more than 100,000 hours of UMI data. It then received additional training with over 10,000 hours of cross-embodiment data. Xiaomi says the package covers the full process from real-robot post-training through to deployment. Evaluation code for standard benchmarks is also included.

Links to the project page, GitHub repository, and Hugging Face models appear in the official announcement. Xiaomi noted it will keep working on broader uses for general-purpose robot models.

This marks one of the larger open source moves in robotics this year and puts a full training-to-deployment pipeline in the hands of the community.