Merrin started as a simple trade. Talk about your day, your plans, your doubts, and find a journal waiting for you afterward. No blank page. No pressure to write something that sounds meaningful.

That was the whole pitch, and it survived about a week of building. Then came the second requirement: those conversations should stay on your phone. Suddenly a small journaling app had a large engineering problem, and most of what we learned solving it is not in any feature list.

This is the build story of that work. It is not a launch announcement, and it is not a polished retrospective. It is the record of the decisions that shaped the app, including the ones we reversed. Tap through the parts below; the choices are the story.

It was supposed to be a journal

The first version was a conversation, nothing more. You talk, and the useful parts are saved and made findable later. The point was to remove the blank page, because anyone who has tried to keep a journal knows what happens to the blank page.

The idea we kept circling back to was this:

What if an AI could be useful because it remembered your life, not because it knew everything on the internet?

Answering that question honestly meant on-device work from the beginning: memory, search, voice, and the model itself. Everything else followed from there.

The first decision: the easy way we didn't take

Using a cloud AI API would have been the easy way. Faster to prototype, fewer runtime surprises, and access to hosted models far larger than anything a phone can hold. For a journaling app, that is the normal architecture.

We chose on-device models instead, because the promise only holds if conversations do not leave the device. That decision set the difficulty of everything that came after it. Try the trade below with the question we actually asked ourselves.

Explore the tradeoff

Which version would you have built?

What gets easier
Core conversations can run after model installation, without a cloud inference request.
What gets harder
Model compatibility, downloads, memory limits, heat, latency, and recovery become your problem.

What we chose: on-device. The core promise mattered more than the convenient architecture.

Same decision, two different bills to pay. We picked the one where the cost lands on us instead of on the conversation.

Design decisions are usually presented as a list of wins. This one was a trade, and we wrote it down that way. The convenient architecture loses to the one that keeps the promise, and then you spend months paying for the choice in downloads, memory budgets, and heat.

Our phones became model laboratories

We began with small open models, including Gemma and Qwen variants. The early question was shallow: can this run at all? That quickly turned into a harder one: can this hold a good conversation without crashing, stalling, or losing its place?

Model size and benchmark scores stopped being enough. The only test that mattered was behavior inside Merrin. Tap through the shipped tiers to see the trade each one makes.

Tap a model to inspect the tradeoff
Language model

2.59 GB

Default · the practical starting point

The recommended model, chosen to lower the initial download and device burden.

Cataloged model files, decimal GB. Download size is not a measure of intelligence or runtime speed.

The most useful discovery was about context windows. A model's advertised context length was not the same as the context that ran reliably in the mobile runtime. The bundles the app ships are pinned to a tested, runtime-safe window of 4,096 tokens, and retrieval is designed around that budget instead of a spec sheet.

A lovely voice is useless if it keeps you waiting

Speech was supposed to be a feature. In practice it was another system to tune: recognizing words, understanding the result, and speaking a reply without killing the rhythm of the conversation.

One candidate, Matcha-TTS, did not survive the benchmarks. Against the Inflect-based voice already working in the app, it was far too slow on the same test setup. The recorded difference was 12 to 17 times slower, and it was removed from the curated list on September 30.

12–17×
Slower. Same benchmark setup. That is how Matcha-TTS compared with the shipped voice in the recorded test. The next voice on the list was not the prettiest one. It was the one that answered before you stopped waiting.
We stopped asking which model sounded more impressive and started asking which one people would actually wait for.

A voice can be beautiful and still be wrong for a conversation. If the reply arrives after the moment has passed, the quality of the audio does not rescue it.

It saved the memory and still acted like it forgot

This one was more frustrating than any slow model. Merrin could keep conversation history and still fail to bring the right information into a reply.

Saving something is one step. The model has to be shown the right thing, at the right moment, in a context window that can actually fit it. So we stopped treating memory as one magical feature and traced the path instead. Follow a memory through the five stages.

Follow a memory through the app
Capture: did Merrin understand what was worth keeping? A useful fact must first be recognized inside an ordinary conversation.

These were investigation points, not five confirmed bugs. Tap each stage to see why an apparently saved memory may not reach a reply.

Fixing context limits and background job handling helped. The bigger shift was architectural: conversations, persistent memories, journal entries, and derived insights needed clear boundaries. The engineering deep dive covers how that memory loop works under the hood.

Yes, we changed the chat bubbles. Then changed them back.

At one point Merrin's replies were open text. Then we tried mirrored chat bubbles. Then we returned to open text. Each version solved a problem and created a different one. Switch between the iterations.

Chat design iterations
I keep going back and forth on this decision.
You came back to it twice this week. What changed since the last time?
Current direction: a quiet, open Merrin reply. Your message stays contained, and the hierarchy feels less like two people trading identical cards.

Conceptual recreations of the iterations, not screenshots from historical builds.

It was not a grand rebrand. It was the ordinary, slightly maddening work of getting a screen to feel right.

Navigation changed too. A split tab bar looked clever on paper but isolated Merrin at the edge of the interface, so it was dropped. Tapping into Merrin also stopped auto-opening the microphone. Talking should be a deliberate action, not something the app assumes.

It works. Why would you open it tomorrow?

By this point we had spent a lot of time making Merrin possible. On-device inference ran. Memory worked, most of the time. The voice was quick enough. And we had not spent nearly enough time on the question a person asks before opening an app for the tenth day in a row:

It works. Why would I open it tomorrow?

Keeping a record of past conversations is useful. But asking someone to open an app just to browse old conversations is not much of a habit. A history you never revisit is a storage feature, not a reason to return.

A better reason to return

The more interesting possibility is making the history tell you something you had not noticed. That question pushed the product toward Now, Signals, and Rewind: three different ways to see your own thoughts over time, not another analytics dashboard. Explore them below.

The direction that came out of the question

What is still on your mind?

Bring forward a recent unfinished thought, rather than asking you to start over every time.

Illustrative Now card

“You were still weighing that launch decision yesterday.”

Proposed experience examples, not claims about insights already generated from a real user's private data.

All three share one rule: nothing is surfaced unless the history supports it. Insights have to be earned from actual entries, not guessed at, which is harder than it sounds and is where we spend a lot of our time.

What we are still working on

The difficult part was never getting AI onto a phone. It was making it quick enough to talk to, dependable enough to remember, and useful enough to return to.

We are still working on the last part. Better retrieval matters. Insights have to be supported by real history, not plausible-sounding guesses. And the honest limits stay in view: on-device models are weaker than frontier cloud models on knowledge-heavy questions, and a local journal is only as safe as the device it is on, so keep your own backup. Our privacy policy describes where the boundary between local and online sits. If Merrin cannot make a person's own thoughts more useful to them, the rest is just an interesting technical demo.

We started out trying to build a journal that writes itself. We are building one that helps you see yourself more clearly instead. It is not finished. That is where the story is.

Frequently asked questions

Why did Merrin choose on-device AI instead of a cloud API?

The promise was that conversations stay with you. A cloud API would have been faster to prototype and would have given access to larger models, but it would mean sending personal conversations to a service. On-device processing keeps that promise true offline and without a server copy.

Why doesn't Merrin let you choose a model?

Model choice is treated as an implementation detail. The app manages downloads, memory budgets, and fallbacks so the person stays in the conversation instead of a settings screen. The engineering deep dive explains how that works.

Why was the Matcha-TTS voice removed from Merrin?

In the same benchmark setup, it was 12 to 17 times slower than the Inflect-based voice already in use. A reply that arrives late breaks the rhythm of a conversation, so it was removed from the curated model list on September 30, 2026.

What does it mean when Merrin remembers something?

Remembering is a chain: capture, save, index, find, and use in the model's context. A memory is only useful when the whole chain works, which is why the app links memories back to the messages that created them and lets you inspect, edit, or delete what it keeps.

What is next for Merrin?

The direction is Now, Signals, and Rewind: surfacing a recent unfinished thought, recurring themes that more than one entry supports, and how your thinking changed over time. The rule for all three is the same: only surface what the actual history supports.