Most AI companions forget you the moment a request ends. The service holds your context in a server-side profile, bills you per token to reload it, and keeps a copy of your private life as a side effect. We built Merrin the other way. Our private AI companion keeps the language model, the memory, and the transcript on the device, and the product is shaped around the limits that on-device AI creates. This post walks through the architecture we landed on, the memory loop that makes a small local model feel like it knows you, and the places where local-first is genuinely worse than the cloud.
If you are building anything that touches personal data, the useful question is not “can we run this locally?” It is “what breaks when we do, and is the trade worth it?” For an AI journal you keep for years, it is. Here is the version of that argument we wish we had read before we started.
Why local-first is an architecture, not a privacy badge
Local-first is often sold as a checkbox. In practice it rewrites nearly every technical decision: which model you can load, how you store memory, how you retrieve it, how you stream tokens to a screen, and how you handle the long tail of old devices that cannot hold a 4 GB weight file in RAM. A privacy policy can promise that data stays on device; only the architecture can keep that promise true when a user is offline, when a phone is thermal-throttling, or when a retrieval query would otherwise be one API call away. The constraint is real, and it is the product.
The upside is not only privacy. A local model is available on a plane, in a hospital, and in a country where your data-protection lawyer has opinions about cross-border transfer. There is no inference bill that scales with how much a person writes, so we can encourage journaling instead of discouraging it. There is no network round trip between a thought and an answer, which matters more than raw tokens per second for the felt quality of a conversation. None of that removes the trade-offs; it just puts them somewhere we control. The privacy policy describes exactly where the boundary sits.
The four pipelines that make a local companion feel alive
A local-first AI companion is not one model. It is a small set of pipelines that each have their own budget. The first is capture: text and speech recognition turn a thought into a message. The second is memory: an embedding model turns each message into a vector, a local index stores it, and a memory store keeps the structured facts worth carrying forward. The third is retrieval: a query is embedded, matched against the index, reranked, and trimmed to fit the context window of the local model. The fourth is generation: the local LLM streams an answer, and optional speech synthesis reads it back. Everything except model download and purchase sits inside the device trust boundary.
The important detail is that those network calls are optional and user-triggered. A model download is a one-time fetch; an entitlement check is a purchase event. Neither carries the content of a conversation. That separation is what lets us say core functionality works offline after the required models are installed and mean it, rather than treating offline as a degraded mode.
The memory loop is a budget, not a database
The naive way to give a model memory is to stuff the whole history into the prompt. It works in a demo and collapses in production: the context window fills, tokens per second drops as the prompt grows, and the model starts ignoring the middle of a long conversation. The other naive option is to store nothing and summarize aggressively, which loses the specific details people actually ask about, such as what they said about a launch. We chose a hybrid. Recent turns stay verbatim inside a fixed window; older turns are embedded for semantic search and distilled into structured memories with a type, a source link, and a confidence score.
That design turns memory into a budget you spend per query rather than a pile you keep forever. Retrieval returns a small candidate set, a reranker orders it, and the prompt builder truncates to fit the model's window with room left for the answer. The chart below shows the shape we design for: memory grows, but the useful recall rate flattens instead of keeping pace, because the reranker is doing the work that raw size used to do.
Retrieval is where a small model earns its keep
A 3B parameter model cannot know the same things a frontier model knows, so the retrieval layer has to be unusually good. Our flow is boring on purpose. The query is embedded with the same model that embedded the corpus, so the vectors live in one space. The vector index returns the top-k nearest neighbors by cosine similarity. A lightweight reranker scores each candidate against the query and against recency, and the prompt builder packs the winners into the system context with their source identifiers attached. Attaching sources is not only for citation: it lets the UI show the user the exact journal entry an answer came from, which is the feature that makes local memory feel trustworthy rather than magic.
The failure mode to watch for is duplicate memories. Early on we embedded every message, so a single idea repeated across a week produced a dozen near-identical vectors that crowded out other results. Deduplicating by cosine distance before insertion, and merging repeated facts into one memory with an updated timestamp, fixed more retrieval quality than any model upgrade we tried.
Latency is a budget you spend in stages
Users do not experience inference latency; they experience the gap between finishing a sentence and seeing a response begin. That gap is the sum of several stages, and the stages have very different costs offline. Speech recognition and the first generated token dominate; embedding and vector search are almost free by comparison. This is why on-device AI feels competitive for conversation even when its raw tokens per second is lower than a cloud model. The cloud pays a fixed network tax before it generates anything, and for short, personal exchanges that tax is most of the wait.
The engineering consequence is that we optimize time to first token much harder than total generation time. Streaming the answer matters, but streaming a slow first token does not. We preload the model during idle moments, keep the embedding model resident, and keep the vector index memory-mapped so a query does not wait on disk. On supported devices that keeps the felt gap small enough that the local model wins the comparison the user actually makes.
On-device and cloud are not one score
It is tempting to ask whether local-first AI is better. They trade across different axes, and a single benchmark hides that. Local wins on privacy, offline reliability, and cost per interaction. Cloud wins on knowledge breadth and, on older hardware, on raw speed. Customization is closer than people expect: a local setup lets you choose the model, the persona, and the retrieval policy, while a hosted API locks those behind a vendor's roadmap. Draw the comparison honestly and the answer stops being ideological.
We deliberately do not expose model names, context-window sizes, or quantization levels in the product. Those are implementation details; the user is deciding whether to trust a companion with their thoughts, not choosing a runtime. The comparison above lives here, in an engineering post, because that is the audience that needs it.
Model and device trade-offs decide your support matrix
The hardest product constraint is hardware. A model that runs in 3 GB of RAM on a flagship phone will swap or crash on an entry-level device, and the difference between 6 and 40 tokens per second is the difference between a conversation and a stall. We sized the shipped model for the middle of the market and made quality adaptive: larger models download on capable devices, and the app falls back to a smaller model when memory pressure spikes. That fallback has to be invisible. If it changes the personality, users notice immediately.
If you are choosing a stack for a local AI product and this shape looks familiar, the trade-offs are similar whether you are shipping a consumer companion or an on-premise assistant for a regulated team.
What we tried that did not work
We shipped and then removed three things, and the reasons are more useful than the final architecture. Full-history prompting was the first casualty. It looked wonderful for the first week of a user's history and then degraded, because the model attended to the most recent and most recent-sounding text regardless of relevance. Per-message embeddings without deduplication was the second. Retrieval got slower and less diverse as history grew, which is the opposite of what the growth curve suggests. Running a single large model on every device was the third. It was excellent on the test fleet and unusable on the long tail, and supporting that long tail is the whole point of a consumer product.
The pattern is the same each time: a technique that wins on the happy path loses on the distribution. Build for the device you do not own, the conversation with three hundred prior turns, and the user who journals every day for a year. The demo is the easy part.
Known limitations and open questions
Local-first is not free. Small models are weaker at knowledge-heavy questions, and we route those honestly rather than pretending otherwise. Thermal throttling means performance is not constant, and a benchmark run on a cool phone overstates what a user sees after twenty minutes of voice interaction. Storage is a real constraint on entry-level iOS and Android devices with 64 GB of total capacity. And memory is genuinely hard to get right: a system that remembers too much becomes noisy, while one that remembers too little stops feeling personal. We have a strong default, not a solved problem.
There is also an honest security caveat. On-device data is only as safe as the device, so we rely on operating-system application isolation and platform secure storage rather than inventing our own. If you lose the phone, you lose the data; that is a feature for some people and a risk for others, and Merrin says so plainly. Our privacy policy describes the boundary between what stays local and the limited network calls, such as model downloads and purchases, that do not.
How to try this approach
If you want to build in this space, start with the retrieval loop before you touch the model. Embed a small corpus, measure recall at different k values, and make the reranker good enough that you can drop the context window without losing answer quality. Only then pick a model, and pick it for your worst supported device rather than your best. Keep every network call user-visible and optional, and attach a source to every memory-backed claim so the user can audit the system.
Merrin is available on supported iOS and Android devices, and you can see how the memory controls work on the home page. If you are designing a private AI companion for a regulated or privacy-sensitive team and want to compare notes, talk to us. We have made most of these mistakes already.