
Your AI, your device,
your silicon.
Chat with Llama, Gemma 4, DeepSeek, Mistral and 100+ more — accelerated by your phone's NPU, your Mac's MLX engine, or any GGUF runtime. Completely offline. Completely free.
Mac: v1.3.7 · 142 MB · macOS 14+ · signed & notarized

What's new in v1.3.7
Platform
Mac app
Signed, notarized universal DMG. macOS 14+. Updates itself.
Agents
Agents that browse
25 skills. Web search, page reading, and your knowledge bases.
Memory
AI Memory
Remembers facts across chats. /remember, /recall, /forget.
Models
MedGemma + Qwen 3.5
On-device medical reasoning and near-Sonnet quality at 4B.
Why FluentAI?
The privacy-first AI agent platform that puts you in control
Privacy First
Your conversations never leave your device. No data collection, no tracking, no cloud required.
100+ AI Models
Run Llama, Gemma, DeepSeek, Mistral locally or connect to Claude, GPT-4, Gemini via cloud.
Voice Chat
Talk to AI naturally with 5 conversation modes — Normal, Interview, Learning, Storytelling, and Translation.
Completely Free
No $20/month subscriptions. Use powerful local models at zero cost, forever.
Knowledge Bases
Upload PDFs and documents to chat with your own data. On-device RAG with semantic search.
Tool Calling & MCP
Built-in tools for search, math, weather, and memory. Connect to GitHub, Slack, Notion via MCP.
Chat Organization
Folders, tags, pinning, branching, and search. Keep your conversations organized your way.
Bring Your Own Model
Import any GGUF model or load directly from Hugging Face. Use any model you want — total freedom.
Multi-Runtime EngineNEW
Same chat, three backends: GGUF, LiteRT, MLX. App picks the fastest one per device automatically.
NPU AccelerationNEW
Snapdragon NPU via QNN delegate. 2–4× faster local inference on supported phones, lower battery drain.
Apple Silicon MLXNEW
Native Metal-backed inference on Apple Silicon Macs. No Rosetta. No fallback.
OpenAI-Compatible ServersNEW
Point at LM Studio, vLLM, LocalAI, Jan or any /v1/chat/completions endpoint. Models auto-discover.
On-Device AI AgentsNEW
Give it a goal, not a prompt. On-device agents plan, search the web, read your documents, and show every step they took. 25 skills built in.
AI MemoryNEW
FluentAI remembers facts across conversations, learning quietly as you chat. Steer it with /remember, /recall and /forget.
Local API ServerNEW
Turn FluentAI into an OpenAI-compatible server on your own network. Point Cursor, Continue, or any /v1/chat/completions client at your device.
On-device VisionNEW
Attach a photo and ask about it. Vision-capable models process the image locally — nothing is uploaded.
Powerful Capabilities
More than just a chat app — FluentAI is a complete AI toolkit

Chat With Your Documents
Upload PDFs, text files, and documents to create knowledge bases. FluentAI uses RAG (Retrieval-Augmented Generation) to search and answer questions from your files — all processed on-device.

Built-in Tools & MCP
FluentAI comes with built-in tools — calculator, web search, weather, date/time, and AI memory. Plus full Model Context Protocol (MCP) support to connect to GitHub, Slack, Notion, and 20+ other services.

Rich Content & Code
Beautiful syntax-highlighted code blocks, LaTeX math rendering, HTML/SVG previews, and full Markdown support. Perfect for developers, students, and researchers.

Templates & AI Personas
Choose from built-in prompt templates or create your own. Set up custom AI personas with unique system prompts — from a coding assistant to a creative writing partner.

Android and Mac, today
Available now on Google Play and as a signed, notarized Mac app that updates itself. The Mac build adds code execution, background inference, and native MLX acceleration on Apple Silicon. Windows and Linux are in development.
Beyond chat
Agents that run on your device
Give FluentAI a goal instead of a prompt. It plans, calls tools, and works through the task — on your own hardware.
Runs on your own model
The agent loop executes against the model on your device. No server sees your task, your files, or the pages it reads.
Searches and reads the web
Agents search the web and pull pages down as clean text, then cite every source they used in the result.
Queries your knowledge bases
Agents search documents you've imported before reaching for the web, so answers stay grounded in your own material.
Shows its work
A colour-coded execution trace records every plan, tool call and response. Continue in chat to keep going with the full model.
25 skills built in
Invoke any of them with the /agent command.
Three agent runs per day are free. Scheduled agents are a Premium feature.
Inference engines
One app. Three inference engines.
FluentAI automatically picks the fastest runtime for your device — GGUF for universal support, LiteRT for Android NPU/GPU acceleration, MLX for Apple Silicon.
On an Apple Silicon Mac, FluentAI runs MLX models natively on the GPU via Metal.
FllamaRuntime
// Runs in every FluentAI build
- →Vulkan on Windows · Metal on Apple · OpenCL on Adreno (Q4_0 and Q5_K)
- →Gemma 4 architecture backport (ISWA dual-cache, MoE 128 experts)
- →KleidiAI v1.23.0 (SME2 + Q4_K paths)
- →KV cache TQ4/TQ3 quantization
LiteRTRuntime
// Android
- →Snapdragon NPU via QNN delegate
- →SoC-aware backend selection: QNN > GPU > CPU
- →Play Feature Delivery — no bloat at install
- →MTP speculative decoding · ~2× faster on Gemma 4n
MlxRuntime
// Mac (Apple Silicon)
- →Real Apple MLX inference on M-series + A17 Pro+
- →Multi-file parallel download from Hugging Face
- →Metal-native — no Rosetta, no fallback
- →1-bit quantization: 7B model in ~1.75 GB
Works with your favorite models
Run models locally on your device or connect to cloud providers — your choice
Llama
On-deviceGemma 4 E2B / E4B
Google · Apache 2.0
On-deviceDeepSeek
On-deviceMistral
On-devicePhi
On-deviceQwen 3.5
Alibaba
On-deviceNemotron
NVIDIA
On-deviceMedGemma
Medical AI
On-deviceEmbeddingGemma
On-device RAG
On-deviceClaude
Anthropic
CloudGPT-4
OpenAI
CloudGemini
OpenRouter
200+ models
CloudLM Studio · vLLM
LocalAI · Jan · /v1
OpenAI-compatOllama
Local server
InfrastructureSee it in action
A beautiful, intuitive interface on your phone and your Mac
Your data stays on your device
FluentAI is built from the ground up with privacy as the foundation, not an afterthought
Zero Data Collection
No telemetry, no tracking, no analytics. Your conversations are yours alone.
Offline Capable
Run AI models entirely on your device. No internet connection needed.
Open Models
Runs open-weight models from Google, Meta, Alibaba, DeepSeek and NVIDIA. Import any GGUF file. No vendor lock-in.
How FluentAI compares
Hardware acceleration, BYO local servers, and total model freedom — the differentiators cloud apps can't match.
| Feature | FluentAI | ChatGPT | Claude | Gemini |
|---|---|---|---|---|
| Price | Free (local models) | Free / $20/mo | Free / $20/mo | Free / $20/mo |
| Privacy | On-device, zero collection | Cloud, data used for training | Cloud-based | Cloud, data used for training |
| Offline Mode | ✓ | ✗ | ✗ | ✗ |
| Model Choice | 100+ models | GPT-4 only | Claude only | Gemini only |
| Hardware AccelerationNEW | NPU + GPU + Metal + CPU | Cloud only | Cloud only | Cloud only |
| BYO Local ServerNEW | LM Studio · vLLM · LocalAI · Jan · Ollama | ✗ | ✗ | ✗ |
| BYO Model (GGUF / HF)NEW | ✓ | ✗ | ✗ | ✗ |
| Voice Chat | ✓ | Paid | Paid | ✓ |
| On-device Agents | ✓ | ✗ | ✗ | ✗ |
Frequently Asked Questions
Everything you need to know about FluentAI





