
Your AI, your device,
your silicon.
Chat with Llama, Gemma 4, DeepSeek, Mistral and 100+ more — accelerated by your phone's NPU, your Mac's MLX engine, or any GGUF runtime. Completely offline. Completely free.
Mac: v1.3.7 · 142 MB · macOS 14+ · signed & notarized

What's new in v1.3.8
On Android today. The Mac app is on v1.3.7 and picks these up with its next release.
Characters
Character chat & calls
14 built-in characters, each with its own voice. Tap call and talk out loud.
Images
Image Studio
Generate images entirely on-device. No cloud, no upload, nothing leaves your device.
Models
The 2026 edge catalog
LFM2.5, Qwen 3.6, Granite 4.0 Nano, Nanbeige 4.2 and Ornith 1.0 join the catalog.
Speed
Speculative decoding
A small draft model races ahead of the big one. Same output, fewer seconds.
Why FluentAI?
The privacy-first AI agent platform that puts you in control
Privacy First
Your conversations never leave your device. No data collection, no tracking, no cloud required.
100+ AI Models
Run Llama, Gemma, DeepSeek, Mistral locally or connect to Claude, GPT-4, Gemini via cloud.
Voice Chat
Talk to AI naturally with 5 conversation modes — Normal, Interview, Learning, Storytelling, and Translation.
Completely Free
No $20/month subscriptions. Use powerful local models at zero cost, forever.
Knowledge Bases
Upload PDFs and documents to chat with your own data. On-device RAG with semantic search.
Tool Calling & MCP
Built-in tools for search, math, weather, and memory. Connect to GitHub, Slack, Notion via MCP.
Chat Organization
Folders, tags, pinning, branching, and search. Keep your conversations organized your way.
Bring Your Own Model
Import any GGUF model or load directly from Hugging Face. Use any model you want — total freedom.
Multi-Runtime Engine
Same chat, three backends: GGUF, LiteRT, MLX. App picks the fastest one per device automatically.
NPU Acceleration
Snapdragon NPU via QNN delegate. 2–4× faster local inference on supported phones, lower battery drain.
Apple Silicon MLX
Native Metal-backed inference on Apple Silicon Macs. No Rosetta. No fallback.
OpenAI-Compatible Servers
Point at LM Studio, vLLM, LocalAI, Jan or any /v1/chat/completions endpoint. Models auto-discover.
On-Device AI Agents
Give it a goal, not a prompt. On-device agents plan, search the web, read your documents, and show every step they took. 25 skills built in.
AI Memory
FluentAI remembers facts across conversations, learning quietly as you chat. Steer it with /remember, /recall and /forget.
Local API Server
Turn FluentAI into an OpenAI-compatible server on your own network. Point Cursor, Continue, or any /v1/chat/completions client at your device.
On-device Vision
Attach a photo and ask about it. Vision-capable models process the image locally — nothing is uploaded.
Character ChatNEW
14 built-in characters, each with their own voice, personality and memory of your conversations together. Three are free; Premium unlocks the rest and lets you build your own. On Android today.
Call Your CharactersNEW
Tap the call button in a character chat and talk out loud. Each character greets you first and answers in its own voice and accent. On Android today.
Image StudioNEW
Generate images entirely on your device — no cloud, no upload, nothing leaves it. Free accounts get 2 images a day; Premium removes the limit. On Android today.
Speculative DecodingNEW
A small draft model runs ahead and the main model verifies it in one pass. Faster GGUF generation on capable devices, with identical output.
18 LanguagesNEW
The whole app is translated into 18 languages — Arabic, Bengali, Chinese, French, German, Hindi, Japanese, Spanish and more. Pick yours in Settings.
Support When It MattersNEW
If a conversation turns toward self-harm, FluentAI surfaces real crisis resources for your region — detected on-device, like everything else.
Powerful Capabilities
More than just a chat app — FluentAI is a complete AI toolkit

Chat With Your Documents
Upload PDFs, text files, and documents to create knowledge bases. FluentAI uses RAG (Retrieval-Augmented Generation) to search and answer questions from your files — all processed on-device.

Built-in Tools & MCP
FluentAI comes with built-in tools — calculator, web search, weather, date/time, and AI memory. Plus full Model Context Protocol (MCP) support to connect to GitHub, Slack, Notion, and 20+ other services.

Rich Content & Code
Beautiful syntax-highlighted code blocks, LaTeX math rendering, HTML/SVG previews, and full Markdown support. Perfect for developers, students, and researchers.

Templates & AI Personas
Choose from built-in prompt templates or create your own. Set up custom AI personas with unique system prompts — from a coding assistant to a creative writing partner.

Android and Mac, today
Available now on Google Play and as a signed, notarized Mac app that updates itself. The Mac build adds code execution, background inference, and native MLX acceleration on Apple Silicon. Windows and Linux are in development.
Beyond chat
Agents that run on your device
Give FluentAI a goal instead of a prompt. It plans, calls tools, and works through the task — on your own hardware.
Runs on your own model
The agent loop executes against the model on your device. No server sees your task, your files, or the pages it reads.
Searches and reads the web
Agents search the web and pull pages down as clean text, then cite every source they used in the result.
Queries your knowledge bases
Agents search documents you've imported before reaching for the web, so answers stay grounded in your own material.
Shows its work
A colour-coded execution trace records every plan, tool call and response. Continue in chat to keep going with the full model.
25 skills built in
Invoke any of them with the /agent command.
Three agent runs per day are free. Scheduled agents are a Premium feature.
Inference engines
One app. Three inference engines.
FluentAI automatically picks the fastest runtime for your device — GGUF for universal support, LiteRT for Android NPU/GPU acceleration, MLX for Apple Silicon.
On an Apple Silicon Mac, FluentAI runs MLX models natively on the GPU via Metal.
FllamaRuntime
// Runs in every FluentAI build
- →Vulkan on Windows · Metal on Apple · OpenCL on Adreno (Q4_0 and Q5_K)
- →Gemma 4 architecture backport (ISWA dual-cache, MoE 128 experts)
- →KleidiAI v1.23.0 (SME2 + Q4_K paths)
- →KV cache TQ4/TQ3 quantization
LiteRTRuntime
// Android
- →Snapdragon NPU via QNN delegate
- →SoC-aware backend selection: QNN > GPU > CPU
- →Play Feature Delivery — no bloat at install
- →MTP speculative decoding · ~2× faster on Gemma 4n
MlxRuntime
// Mac (Apple Silicon)
- →Real Apple MLX inference on M-series + A17 Pro+
- →Multi-file parallel download from Hugging Face
- →Metal-native — no Rosetta, no fallback
- →1-bit quantization: 7B model in ~1.75 GB
Works with your favorite models
Run models locally on your device or connect to cloud providers — your choice
Llama
On-deviceGemma 4 E2B / E4B
Google · Apache 2.0
On-deviceQwen 3.6
Alibaba
On-deviceLFM2.5
Liquid AI · edge-first
On-deviceGranite 4.0 Nano
IBM
On-deviceNanbeige 4.2
Nanbeige Lab
On-deviceOrnith 1.0
Open weights
On-deviceNemotron 3 Nano
NVIDIA
On-deviceDeepSeek
On-deviceMistral
On-devicePhi
On-deviceMedGemma
Medical AI
On-deviceEmbeddingGemma
On-device RAG
On-deviceClaude
Anthropic
CloudGPT-4
OpenAI
CloudGemini
OpenRouter
200+ models
CloudLM Studio · vLLM
LocalAI · Jan · /v1
OpenAI-compatOllama
Local server
InfrastructureSee it in action
A beautiful, intuitive interface on your phone and your Mac
Your data stays on your device
FluentAI is built from the ground up with privacy as the foundation, not an afterthought
Zero Data Collection*
Your conversations are yours alone — never uploaded, never used for training.
Offline Capable
Run AI models entirely on your device. No internet connection needed.
Open Models
Runs open-weight models from Google, Meta, Alibaba, DeepSeek and NVIDIA. Import any GGUF file. No vendor lock-in.
* When you run a model on your device, your prompts, chats, documents and model output never leave it. Cloud providers, connected servers and the free starter chat are opt-in, and send your messages to that service. Separately, the app collects usage and crash telemetry — linked to your account if you sign in — plus an advertising identifier on the free tier. See exactly what we collect.
How FluentAI compares
Hardware acceleration, BYO local servers, and total model freedom — the differentiators cloud apps can't match.
| Feature | FluentAI | ChatGPT | Claude | Gemini |
|---|---|---|---|---|
| Price | Free (local models) | Free / $20/mo | Free / $20/mo | Free / $20/mo |
| Privacy | On-device, zero collection | Cloud, data used for training | Cloud-based | Cloud, data used for training |
| Offline Mode | ✓ | ✗ | ✗ | ✗ |
| Model Choice | 100+ models | GPT-4 only | Claude only | Gemini only |
| Hardware Acceleration | NPU + GPU + Metal + CPU | Cloud only | Cloud only | Cloud only |
| BYO Local Server | LM Studio · vLLM · LocalAI · Jan · Ollama | ✗ | ✗ | ✗ |
| BYO Model (GGUF / HF) | ✓ | ✗ | ✗ | ✗ |
| Voice Chat | ✓ | Paid | Paid | ✓ |
| On-device Agents | ✓ | ✗ | ✗ | ✗ |
| On-device Image GenerationNEW | 2/day free, unlimited on Premium | Cloud only | ✗ | Cloud only |
| Characters & Voice CallsNEW | 14 built-in, plus your own | Cloud only | ✗ | Cloud only |
Frequently Asked Questions
Everything you need to know about FluentAI





