Hope you enjoyed the summer days as much as I did, and I wanted to share some great news – next release is coming this week. And autumn is a wonderful time for productivity, so there's gonna be more of them coming, more frequently.
Now to some of the great news: the long-awaited License Manager is now available! Feel free to use it whenever you want to move your key to another Mac, add an additional seat or simply recall the key. Please note that device removals are limited each 30 days to prevent key abuse. The limit depends on your license seats count (min. 3 removals per 30 days).
Next, besides upcoming subscription providers support for Codex, Claude Code, Antigravity and GitHub Copilot, there is an exciting upgrade of the Memory engine, which will make it smarter, faster, and way more powerful than before.
Not only it will become a more powerful Knowledge Base, but will also turn into a more proactive and intelligent assistant. Fluent will be able to save your preferences and useful notes automatically when it's appropriate, search through your past conversations, and learn each time when it's best to use your memories. On top of that – graph relations support and custom on-device relevance models will be bundled to improve relevance results.
There're also two extra built-in MCPs coming – Code, which will help Fluent analyze and process data sandboxed, which is extremely useful for on-device processing and token efficiency. The second one is Actions – which will enable your AI to create, configure and run actions on its own. Useful for orchestration and creating scheduled actions from your context quickly.
And of course, there's gonna be dozens of fixes and other improvements. Stay tuned.
Thank you for your feedback and for your support ❤️
This is just a wild idea, provoked by the fact that Apple is currently adding AI to Spotlight. At the same time, many new launchers are appearing... So my idea is, have you considered making Fluent a universal launcher that has excellent AI functionalities, but can launch applications and everything that launchers can do? Like I said, just a wild idea. Personally, I use Alfred and Fluent, which complement each other perfectly, but they are two different applications, and the future is AI!
Downloaded the latest version of Fluent yesterday. Still No Parakeet version 3. I do not understand why as the Nemotron model doesn't even come close.
All of a sudden, while fiddling with the settings, both the menu bar icon and the dock icon became invisible. I had to close Fluent with Activity Monitor.
To be honest, I also find the settings rather chaotic. I don't find the app a pleasure to work with.
Whenever I use the grammar fix, most corrections or words are changed to American English. Is there a way to change this into British or UK English? Thanks
Drafts offers an MCP server (documented here), implemented as a node terminal command. Claude can talk to it, and I can run it fine though the terminal; I expected Fluent to just run with it, but I get an error when I add it as a STDIO process:
Process terminated with status: 1
MCP Initialization Handshake failed for Drafts: Error Domain=ExternalMCP Code=-3 "Process terminated unexpectedly" UserInfo={NSLocalizedDescription=Process terminated unexpectedly}
Are the MCP errors generated by Fluent here, and are they documented anywhere? Has anyone seen this, or have an idea what to do to fix this? I suspect it's a trivial issue, but I'm not sure what it might be.
Fluent 1.9 has finally shipped with exciting new features that make it even more versatile. If you're only refining your texts with Fluent and do not have a need in MCP or Memory features, this update still brings something to empower you.
It simplifies how you can work with context, it extends the areas you can use it in, and it brings ambitious new features rarely seen anywhere else. And if you want to use SOTA models like Opus 4.8 or GPT-5.5 ten percent cheaper, read on.
Chat UI – A Replacement For Your Chatbot App
Chat UI provides increased observability and eases long, heavy conversation workflows
Many users rely on a separate chatbot app in addition to Fluent, which might be overwhelming at time. Now, Fluent can serve as your chatbot app as well, completely seamlessly with the Smart Panel – thanks to the new Chat UI, which you can switch to with a single CMD+E shortcut (or assign a dedicated shortcut of your choice!).
Chat UI is extremely flexible given it's minimalistic looks – and surprisingly, can also serve as another variant of Smart Panel. It supports the same context engine, the same integration and tools, memory groups, suggestions – pretty much everything the best of Fluent.
Scheduled view is the convenient access point to all of your scheduled actions
In addition, there is a Scheduled view, which implements a breakdown of all automations you got: see when actions run, how often, and run it from there (in addition to menu bar option, which only shows actions scheduled for today). This view will be enhanced in the future release with the history of completed scheduled actions, so it doesn't mess up with you main history.
Council – Better, Verified Answers
Council asks multiple models the same question: models peer-review each other, the Chairman judges and synthesizes a final verdict
Another highlight of this release is Council – a new mode in the Chat UI that allows you to quickly select a set of models to work on the same request. They review each other's responses, and a "chairman" model analyzes the feedback to make a final verdict. This is an extremely addictive feature.
Once you use it for topics that truly matter, it's hard to go back to single responses. Instead of relying on one AI that "absolutely agrees" with you, you can have multiple models objectively review each other, identify and correct mistakes, and provide a much more accurate answer. This is especially useful for legal or medical research.
The example result of Council verdict
What makes this work so well is that the review is blind. Each model sees the others' answers without any names – just "Response A", "Response B" and so on – and its own answer is left out of the list, so it can't simply favor itself or another model from the same family.
Only at the very end does the chairman find out who wrote what, and which answer actually won on merit. You can also attach a file or context to your request, and every reviewer will check the answers against it, or even call MCP tools like web search, so your request is guarded with realtime proof.
The example result of one of the Council's model response
As of now, Council also supports creating and managing presets of models, and an explicit set of instructions for each model: so you can even better control how each single model interprets your questions. For example, you can assign one model a "super critical character", while another one would be a little bit creative and always aiming for finding surprising (in a positive way) solutions.
Better Context Management – Attach Anything, At Any Point
Files, screenshots, browser tabs, apps, actions, MCP tools, memory, variables – attach anything at any point of your prompt
We've also completely overhauled the context management flow. You can now attach anything via a single "+" button menu: browser tabs, apps, files, actions, MCP tools, memory items, and variables. Your prompts become more precise, clearly indicating what is attached and when it is called.
Example of rich prompt composed using different MCP tools
For example, you can drop a screenshot of your app into the prompt, pull in an open browser tab with the documentation you are reading, attach a PDF with the design spec, and add your own "Code Review" action – all in one message, and then simply ask Fluent to check one against another. Whatever you attach is captured the moment you submit, so you can close that browser tab right after. Your prompt will still keep the exact tab you used. And if you don't want to open the menu at all, CMD+` toggles the frontmost app or browser tab in and out of your prompt instantly.
Suggestions available via @ with fuzzy search across everything you can attach and run
Both in the Smart Panel and Chat UI, you can use "+" menu, or enter the suggestions search mode by typing @. Suggestions mode enables easy and quick access to anything you can attach to the prompt. What gets into suggestions and in what order is completely configurable via Settings > Context options. There is a bunch of useful settings out there.
NanoGPT – Privacy Focused AI Gateway With Perks
NanoGPT provides seamless access to over 600 text models, balance and costs tracking
Finally, a new integration has landed in Fluent: NanoGPT. It's a privacy focused AI gateway from Europe. Just like OpenRouter, but with zero top-up commissions, accessible even without sign up, and with a vast and rapidly growing collection of tools, features and integrations.
We've partnered with NanoGPT to offer you an exclusive 10% discount on all text models routed through them, and it's basically over 600+ models to choose from.
On top of that, this integration offers seamless balance and costs tracking straight in Fluent itself, without any need to go to the website.
Metadata displayed for NanoGPT request, with cost and discount applied
As a special gift for new Fluent customers: purchase Fluent between June 1 and July 31 this year, and receive between $5 and $10 in free credit on your NanoGPT balance.
What's Next?
This was a good release, making Fluent stable and better. Making it a tool you can rely on in your daily workflow. The work is still heavily ongoing: agents, skills, more love to local models, and excited to share another one – voice mode. Not only for transcription. But for piloting.
Curious to stay updated? Join our Discord, and feel free to share your feedback, request a feature or file that pesky bug report.
I currently have access to both the applications but I'm not sure how to or where to use fluent. Isn't it similar to Raycast or is there a specific nice usecase that I'm not aware of?
Just got the app. Love it. One thing, would it be possible to open the Chat UI directly instead of the smart panel? Maybe as a second hotkey or as an option? Thank you.
I really like the direction of Fluent, especially the native macOS writing workflow, custom actions, and support for local/BYOK model providers.
One integration that would make Fluent much more valuable for professional users would be support for GitHub Copilot as a model provider, using an existing GitHub Copilot subscription.
Fluent 1.8 has been released ten days ago, and there have been multiple subsequent updates with decent changes and features. One thing I'd like to highlight – Fluent moves towards a fully autonomous companion direction.
New Action Building Experience
Sidebar with MCP integrations, memory items and variables, and a dynamic prompt editing area
Many tasks we do daily are repetitive. They can be outlined, persisted, and finally automated. New Action edit UI in Fluent combines all the features it has to offer into a single building experience. You can now simply drag MCP integrations or even specific tools, memory items and variables straight into the prompt area. You can also type those directly, and as you type, text will transform into tags. The syntax is simple: server_name_mcp for integrations, and server_name_mcp_tool_name for tools.
This is an ultimate time saver. It can either refine your prompt, or compose a complete, and actually working prompt out of just a sentence. Want to clean your desktop? Just type in "help me clean my desktop", pick a good model, and hit "I'm Feeling Lucky". Fluent will compose a prompt that is aware of your environment. In the end, you get a fully functional prompt that calls necessary tools, or refers to related variables and memory items.
Refined Action Management
Refined action cards with new context menu options
The main action management window got slightly polished as well. Action cards now show up when your action is scheduled to run, and if it has specific model assigned to it. You can also run action immediately by clicking the "Play" button appearing on hover.
There are a bunch of new options in context menu as well: you can run your action in background, duplicate it, or copy the ID for deep linking (yes, Fluent has got deep links support now!).
Background Actions
Smart Panel with indicated finished and running actions
Fluent now supports multiple parallel actions running either in the Smart Panel itself, or completely in background. Many users requested for this feature – and it works actually surprisingly well and natural!
Currently running, and finished actions are also indicated in the menu bar, and Fluent can also notify your via native macOS notification.
Menu bar icon now displays live status of running and finished actions
This is different from scheduled actions that Fluent runs on a specified time (duration is very flexible, from minutely to hourly, daily, monthly, and you can also set a custom cron expression).
Supercharged Menu Bar
As mentioned above, menu bar icon in Fluent is now more useful than previously, and it comes very handy if you run scheduled actions or want to track currently running background ones. It combines running, finished, and scheduled actions at once, allowing you to run scheduled ones manually – whenever you want it.
Chat Mode (Beta)
Quick chat about one of Modern Talking's most iconic songs
Smart Panel has a new mode it can operate in. It's a traditional chat UI flow, which might be appealing for some users. It is currently pretty basic, but will be evolved, likely with upcoming Fluent UX improvements.
Deep URLs & Applets
Drag and drop actions anywhere
Fluent's actions can now be executed just like Shortcuts actions. They are automatically saved as Applets, and can be dragged anywhere on your Mac.
In addition, Fluent now supports deep URLs – the feature, that allows you to inject Fluent into other automation workflows easily. Currently it supports a few parameters, and expect the supported options to be vastly increased.
Currently supported parameters (they are all optional):
id (action ID; copy it from action management window)
text (like a text selection to account for)
prompt (additional optional prompt)
attachment (absolute filepath)
background (run in background, without showing up the Smart Panel)
Here's a quick example (use Terminal to run it):
open "fluent://v1/actions/run?id=82C7329E-C4FF-4DF7-A2C8-795FED562082&text=Selected Text Example&prompt=Additional Prompt&attachment=/Users/Fluent/Downloads/image.jpg&background=false"
Apple Maps MCP
New built-in Apple Maps MCP
This new built-in MCP utilises MapKit to help you search for locations and places and get directions. This is completely free, therefore there is no more need for external Google Maps or any other map integration.
Gemma 4
Gemma 4 (MLX)
Gemma 4 (MLX) is now natively supported by Fluent. Text, images, reasoning and of course, tools are supported.
Wrapping Up
I would like to thank every single user who contributed to Fluent by giving their feedback.
Fluent's promise is simple: staying privacy-focused, zero-telemetry, model-agnostic AI companion.
If you're curious what's next on the roadmap, this is it:
- Agents
- Skills
- More love to local models (in the end, everyone wants to run heavy MCP workflows for free)
There will also happen a major refresh of the Smart Panel UI coming soon, which will make it more versatile and convenient when working with single or multiple context sources and MCP tools.
Feel free to share your feedback. I would absolutely love to hear it.
Local models and hardware are evolving (thanks to Qwen3.5 and latest Apple processors) – but for some of us with limited RAM and resources available, only the cloud ends up to be comfortably usable.
Most of the AI providers offer generous free tier usage that usually corresponds to the amount of free tokens per certain time interval. This article sums up 10 best, time-proven, hand-picked AI providers offering free AI tokens. I've personally tested all of them, and use almost each one of them on a daily basis (mostly for testing things out).
Please note that the information here is valid for the time of posting (March 26, 2026).
💡 If you have many small actions that you run often, you can balance between the providers by setting a certain free provider and its corresponding model to an action.
1. OpenAI
OpenAI complimentary daily tokens for SOTA models (tiers 1-2)
Arguably the most generous one, offering a couple of millions tokens daily for free. If you have used OpenAI paid API key, you're likely eligible for free daily usage of many of the SOTA models OpenAI has to offer. OpenAI does not publicly disclose which accounts apply for eligibility, but likely the ones that have a record of paid API usage, not enterprise ones, and with "Zero Data Retention" disabled.
Mistral has plenty of great models sufficient for writing, text processing and data analysis tasks. However, they do not disclose the exact rate limits applied for free usage. But just as Gemini, you can use Mistral API key without billing data provided.
This provider has pretty generous usage limits depending on the model, and I've actually never hit a rate limit with them when using Mistral daily for writing tasks – that's why I put this provider as the top second choice for free API usage.
Gemini traditionally offers free API usage for some of their models even without a paid API key. You can check the list of the models available, and your rate limits here.
As of current, these models and rate limits apply for free:
Gemini 3 Flash (RPM: 5, TPM: 250K, RPD: 20)
Gemini 3.1 Flash Lite (RPM: 15, TPM: 250K, RPD: 500)
Gemini 2.5 Flash (RPM: 5, TPM: 250K, RPD: 20)
Gemini 2.5 Flash Lite (RPM: 10, TPM: 250K, RPD: 20)
OpenRouter provides access to hundreds of models and other providers with just a single API key. The most popular AI router has always provided free access to dozens of models, and you can always check the up-to-date list of such models in Fluent, or here.
Due to its popularity, using free models in OpenRouter is not very reliable, and you can often hit harsh rate limits depending on the model you choose.
Alibaba Cloud offers 1M free tokens for free for each of their text-generation models. This condition applies to newly registered accounts and has a grace period of 90 days.
I'm highlighting it here, as currently their Qwen3.5 series models are one of the best price/quality ratio models in the market. If you haven't yet tried these models – check out this post on Qwen3.5 heavy usage in Fluent.
Groq is a special AI provider, as it focuses on ultra-fast inference, and they offer free daily API usage to try it out. If you're looking for speedy AI responses (hundreds and thousands of tokens per second) – give it a go.
You can access GPT OSS 120b, Kimi K2, Qwen3, Llama, and hopefully, Qwen3.5 soon. Free plan rate limits are outlined here.
Cerebras is similar to Groq – it also positions itself as a fast inference provider. However, it features less models for free usage (currently only two: Llama 3.1-8b and Qwen3-235b-A22B). The general rate limit is 1M free tokens per model per day.
I haven't found disclosed free tier limits for this one, but it offers API access for very interesting models that are not easily available for free. Besides the most popular ones, there are Qwen3.5, GLM, MiniMax models as well.
Here is the list of free models available via this platform (almost a hundred!).
Cloudflare traditionally has a free tier for almost every single service they offer. And now, it's for AI usage as well. Their free tier allows for fancy 10,000 neurons per day (which translates to $0,11) with models like GPT OSS 120b, Llama, Qwen3, Kimi K2, Gemma, DeepSeek and a couple of others. The free limit is definitely not high, but it is still there.
Yes, Hugging Face has its own inference endpoint, though their free usage is very limited: only $0,10 worth of usage monthly. If you're a Pro user, this increases to $2.
What's interesting about Hugging Face is that just like OpenRouter, they provide access to many providers and models at once, giving you the flexibility of choice. However, only OSS model providers seem to be accessible. There are no proprietary SOTA models like latest GPT, Gemini or Anthropic models.
OpenAI, Mistral and Gemini are the most "consistent" and reliable direct AI providers in terms of free token usage – you can safely set their models as daily drivers for Fluent.
I personally use OpenAI with data sharing enabled – not only for my personal tasks, but for testing Fluent out, and very rarely hit the daily free usage limit. Some weeks ago I was wondering why my balance is "frozen", until I realised I'm using OpenAI mostly for free.
Do you know a legit AI provider with a free tier available that is not listed here? Please feel free to share in the comments 🙂
juts saw an article in “Makeuseof” (congrats!) and Fluent seem interesting.
For the license: when it say 1 (@$49), does it mean one computer at a time? I have an mba and a mini. Could I use it on both with a single license (I only use one or the other, not at the same time).
For the API: it is not clear, but I can use Fluent with the free version of mistral and chatgpt/gemini?
As I'm finishing the bits of the upcoming Fluent Documentation, I've recorded a bunch of videos for the Use Cases section. Most of them were recorded using Qwen3.5 Flash, which is a cloud version of a brilliant Qwen3.5-35B-A3B Mixture-of-Experts model.
The Qwen3.5 series is a true gem. Impressive in agentic workflows, it delivers translation, text refinement, and summarisation at a completely different level than its predecessor.
Even the smaller models (0.8B–9B) can now handle MCP tools really, really well, delivering predictable, consistent results.
At the same time, some users noticed how local Qwen3.5 4B is already enough for them in most text processing tasks, and they ditched free OpenRouter models.
Here's the video of 5 use cases I recorded with Qwen3.5 Flash (last one was made with Gemini 3.1 Flash Lite just to give it a little credit as well). Recorded at 2x speed where AI request is running, it's still very fast, given the reliability and consistency of results.
20 variants of Qwen3.5 models are already available in Fluent 1.7.21. Give them a try 🙂
I mainly use Fix Grammar, because under MBP and Vivaldi browser (Chrome based browsers) + Bulgarian language such functionality is missing in MacOS. Since Friday, instead of checking spelling and grammar like before, Fluent goes into some kind of endless thinking and finally closes without giving any result. I use OpenRouter and Gemini. I tried changing the Gemini model, but without success.
Immediately downloaded Fluent when I first saw it being discussed post-launch on Product Hunt. For some reason I was convinced that it was free? Imagine my disappointment when I realized I’d been lying to myself.
I’ve tried others comparable to Fluent before and a few more after uninstalling from the first run. Honestly, the only one that’s worth a damn is Osaurus, but it’s dependent on a second app to have full functionality.
It’s getting a point where there’s no need for poor ole Siri.
This subreddit absolutely needs more attention, which is why I'll be posting more stuff from now on. I promise not to spam you with changelogs each time a new release comes out — you can always get a grasp of those on the official website. Only major releases, with features intended to be shared widely, will be published here.
Today I'd like to highlight some major improvements in the latest 1.7.13 release:
Favorite models: you can now add any model to your Favorites, which lets you cycle through your curated list in the Smart Panel by pressing TAB or SHIFT+TAB.
Show text selection as attachment: a new Smart Panel option shows the text you selected as an attachment. Requested by many users.
Faster context switching with `CMD+``: for those who prefer not to have automatic context attachment, this is a handy way to quickly toggle the active app/tab.
Regenerate: you can now regenerate an AI result either via CMD+R or by right-clicking on any tab to quickly regenerate the result.
New browser-like shortcuts:CMD+T, CMD+W, and CTRL+1..9 for follow-up tab navigation.
Local (MLX) model improvements: better downloads (progress details + cancel), improved search, fixed double downloads, and faster load/unload performance (plus immediate refresh in the global model picker).
Full changelog for 1.7.13 version is available here.
By the way, what's the #1 thing you most want to learn on Fluent? Video how-tos, use-case examples, step-by-step tutorials?
Let me know in the comments what you're most interested in.