Skip to main content
Privacy-first AI · now with NPU + Apple Silicon acceleration
llama3gemma4deepseekmistral+ qwen · phi · 100s more

Your AI, your device,
your silicon.

Chat with Llama, Gemma 4, DeepSeek, Mistral and 100+ more — accelerated by your phone's NPU, your Mac's MLX engine, or any GGUF runtime. Completely offline. Completely free.

Mac: v1.3.7 · 142 MB · macOS 14+ · signed & notarized

100+AI Models
3RuntimesNEW
NPU+ GPU + CPUNEW
ZeroData Collection
FreeForever
FluentAI chat interface

What's new in v1.3.7

Platform

Mac app

Signed, notarized universal DMG. macOS 14+. Updates itself.

Agents

Agents that browse

25 skills. Web search, page reading, and your knowledge bases.

Memory

AI Memory

Remembers facts across chats. /remember, /recall, /forget.

Models

MedGemma + Qwen 3.5

On-device medical reasoning and near-Sonnet quality at 4B.

See FluentAI in action

Watch a full product walkthrough or a quick 30-second tour

Long-form product demo

Full walkthrough · all v1.3.0 features

YouTube ↗

30-sec Short

Quick tour · share anywhere

YouTube ↗

Why FluentAI?

The privacy-first AI agent platform that puts you in control

Privacy First

Your conversations never leave your device. No data collection, no tracking, no cloud required.

100+ AI Models

Run Llama, Gemma, DeepSeek, Mistral locally or connect to Claude, GPT-4, Gemini via cloud.

Voice Chat

Talk to AI naturally with 5 conversation modes — Normal, Interview, Learning, Storytelling, and Translation.

Completely Free

No $20/month subscriptions. Use powerful local models at zero cost, forever.

Knowledge Bases

Upload PDFs and documents to chat with your own data. On-device RAG with semantic search.

Tool Calling & MCP

Built-in tools for search, math, weather, and memory. Connect to GitHub, Slack, Notion via MCP.

Chat Organization

Folders, tags, pinning, branching, and search. Keep your conversations organized your way.

Bring Your Own Model

Import any GGUF model or load directly from Hugging Face. Use any model you want — total freedom.

Multi-Runtime EngineNEW

Same chat, three backends: GGUF, LiteRT, MLX. App picks the fastest one per device automatically.

NPU AccelerationNEW

Snapdragon NPU via QNN delegate. 2–4× faster local inference on supported phones, lower battery drain.

Apple Silicon MLXNEW

Native Metal-backed inference on Apple Silicon Macs. No Rosetta. No fallback.

OpenAI-Compatible ServersNEW

Point at LM Studio, vLLM, LocalAI, Jan or any /v1/chat/completions endpoint. Models auto-discover.

On-Device AI AgentsNEW

Give it a goal, not a prompt. On-device agents plan, search the web, read your documents, and show every step they took. 25 skills built in.

AI MemoryNEW

FluentAI remembers facts across conversations, learning quietly as you chat. Steer it with /remember, /recall and /forget.

Local API ServerNEW

Turn FluentAI into an OpenAI-compatible server on your own network. Point Cursor, Continue, or any /v1/chat/completions client at your device.

On-device VisionNEW

Attach a photo and ask about it. Vision-capable models process the image locally — nothing is uploaded.

Powerful Capabilities

More than just a chat app — FluentAI is a complete AI toolkit

Chat With Your Documents

Chat With Your Documents

Upload PDFs, text files, and documents to create knowledge bases. FluentAI uses RAG (Retrieval-Augmented Generation) to search and answer questions from your files — all processed on-device.

PDF SupportSemantic SearchOn-device RAG
Built-in Tools & MCP

Built-in Tools & MCP

FluentAI comes with built-in tools — calculator, web search, weather, date/time, and AI memory. Plus full Model Context Protocol (MCP) support to connect to GitHub, Slack, Notion, and 20+ other services.

Tool CallingMCP ProtocolWeb SearchAI Memory
Rich Content & Code

Rich Content & Code

Beautiful syntax-highlighted code blocks, LaTeX math rendering, HTML/SVG previews, and full Markdown support. Perfect for developers, students, and researchers.

Syntax HighlightingLaTeX MathHTML Preview
Templates & AI Personas

Templates & AI Personas

Choose from built-in prompt templates or create your own. Set up custom AI personas with unique system prompts — from a coding assistant to a creative writing partner.

Custom PersonasPrompt TemplatesAuto-fill
Android and Mac, today

Android and Mac, today

Available now on Google Play and as a signed, notarized Mac app that updates itself. The Mac build adds code execution, background inference, and native MLX acceleration on Apple Silicon. Windows and Linux are in development.

AndroidmacOSCode ExecutionAuto-updates

Beyond chat

Agents that run on your device

Give FluentAI a goal instead of a prompt. It plans, calls tools, and works through the task — on your own hardware.

Runs on your own model

The agent loop executes against the model on your device. No server sees your task, your files, or the pages it reads.

Searches and reads the web

Agents search the web and pull pages down as clean text, then cite every source they used in the result.

Queries your knowledge bases

Agents search documents you've imported before reaching for the web, so answers stay grounded in your own material.

Shows its work

A colour-coded execution trace records every plan, tool call and response. Continue in chat to keep going with the full model.

25 skills built in

Invoke any of them with the /agent command.

research-topicmeeting-prepdocument-digesttrip-plannerdaily-journalclipboard-polishquick-textevent-from-clipboardfocus-modeshare-my-location+15 more

Three agent runs per day are free. Scheduled agents are a Premium feature.

Inference engines

One app. Three inference engines.

FluentAI automatically picks the fastest runtime for your device — GGUF for universal support, LiteRT for Android NPU/GPU acceleration, MLX for Apple Silicon.

On an Apple Silicon Mac, FluentAI runs MLX models natively on the GPU via Metal.

FLM

FllamaRuntime

// Runs in every FluentAI build

  • Vulkan on Windows · Metal on Apple · OpenCL on Adreno (Q4_0 and Q5_K)
  • Gemma 4 architecture backport (ISWA dual-cache, MoE 128 experts)
  • KleidiAI v1.23.0 (SME2 + Q4_K paths)
  • KV cache TQ4/TQ3 quantization
LRT

LiteRTRuntime

// Android

  • Snapdragon NPU via QNN delegate
  • SoC-aware backend selection: QNN > GPU > CPU
  • Play Feature Delivery — no bloat at install
  • MTP speculative decoding · ~2× faster on Gemma 4n
MLX

MlxRuntime

// Mac (Apple Silicon)

  • Real Apple MLX inference on M-series + A17 Pro+
  • Multi-file parallel download from Hugging Face
  • Metal-native — no Rosetta, no fallback
  • 1-bit quantization: 7B model in ~1.75 GB

Works with your favorite models

Run models locally on your device or connect to cloud providers — your choice

Llama

Llama

On-device
NEW
Gemma 4 E2B / E4B

Gemma 4 E2B / E4B

Google · Apache 2.0

On-device
DeepSeek

DeepSeek

On-device
Mistral

Mistral

On-device
Phi

Phi

On-device
Qwen 3.5

Qwen 3.5

Alibaba

On-device
Nemotron

Nemotron

NVIDIA

On-device
MedGemma

MedGemma

Medical AI

On-device
EmbeddingGemma

EmbeddingGemma

On-device RAG

On-device
Claude

Claude

Anthropic

Cloud
GPT-4

GPT-4

OpenAI

Cloud
Gemini

Gemini

Google

Cloud
OpenRouter

OpenRouter

200+ models

Cloud
NEW
LM

LM Studio · vLLM

LocalAI · Jan · /v1

OpenAI-compat
Ollama

Ollama

Local server

Infrastructure

See it in action

A beautiful, intuitive interface on your phone and your Mac

FluentAI main chat interface
FluentAI voice chat mode
FluentAI model selection
FluentAI navigation drawer
FluentAI image chat

Your data stays on your device

FluentAI is built from the ground up with privacy as the foundation, not an afterthought

Zero Data Collection

No telemetry, no tracking, no analytics. Your conversations are yours alone.

Offline Capable

Run AI models entirely on your device. No internet connection needed.

Open Models

Runs open-weight models from Google, Meta, Alibaba, DeepSeek and NVIDIA. Import any GGUF file. No vendor lock-in.

How FluentAI compares

Hardware acceleration, BYO local servers, and total model freedom — the differentiators cloud apps can't match.

FeatureFluentAIChatGPTClaudeGemini
PriceFree (local models)Free / $20/moFree / $20/moFree / $20/mo
PrivacyOn-device, zero collectionCloud, data used for trainingCloud-basedCloud, data used for training
Offline Mode
Model Choice100+ modelsGPT-4 onlyClaude onlyGemini only
Hardware AccelerationNEWNPU + GPU + Metal + CPUCloud onlyCloud onlyCloud only
BYO Local ServerNEWLM Studio · vLLM · LocalAI · Jan · Ollama
BYO Model (GGUF / HF)NEW
Voice ChatPaidPaid
On-device Agents

Frequently Asked Questions

Everything you need to know about FluentAI