Ask whether on-device AI beats cloud AI and you get a hundred confident answers that contradict each other. The disagreement has a cause: the two architectures win on different axes, and a single score hides the crossover. A model that fits in a phone will not out-know a frontier model. A frontier model will not answer when the network is down. Both are true at once, and any useful comparison has to hold both.

So we built our own on-device vs cloud AI comparison: six axes, scored, with the rubric behind every number published below. We did not bury the cases where cloud wins. We also build one of the two things being compared, so assume we are biased about the conclusion and judge the method instead. That is what the rubric is for.

Why single scores get on-device vs cloud AI wrong

The question "which is better" assumes the answer is a property of the architecture. It is not. It is a property of the workload sitting on top of the architecture, and small changes to that workload flip the winner.

The winner changes with the workload. A short reflective exchange at midnight, where someone types a paragraph and wants a thoughtful reply: the local model skips the network entirely, and on a supported device it competes on speed while winning on privacy. Summarizing a twenty-page document: the cloud model is faster and more accurate, and no amount of local optimization changes that. The same app on a flight with no connection: one side works, and the other produces a spinner until landing.

Three workloads, three different answers. That is why the scorecard below ranks six axes instead of one, and why you can re-weight them for your own situation.

The rubric behind every score

Every axis is scored from 0 to 100 for each architecture, in five bands:

  • 81 to 100: the axis is a decisive reason to choose this approach.
  • 61 to 80: strong, with no workarounds for most use.
  • 41 to 60: workable, with conditions attached.
  • 21 to 40: weak on its own; you need another system to cover it.
  • 0 to 20: fails at the job.

Every score also carries a basis, because a structural fact, a measurement, and a judgment are not the same kind of claim. Structural means the score follows from where the compute and the data live, and it will hold for any honest implementation. Measured means we measured it, and the conditions sit next to the number. Judgment means it is our read of the trade, with no measurement behind it, and you should treat it as an argument rather than evidence.

The scorecard:

Axis On-device Cloud Basis
Privacy 95 45 Structural: on-device removes the operator-side copy; cloud privacy depends on policy settings
Offline reliability 100 25 Structural: no network dependency once models are installed; cloud has no offline mode
Latency 78 62 Stage-by-stage: local skips the network round trip for short turns; cloud wins long generations and old hardware
Knowledge breadth 55 95 Judgment: phone-sized models know less; retrieval narrows the gap only for personal questions
Cost efficiency 90 70 Usage-dependent: cloud is cheaper at light use, local at sustained heavy use
Customization 82 60 Judgment: local gives model and data control at the price of maintenance; cloud is managed but vendor-led

A score of 100 on offline reliability does not mean the whole product works offline. It means this axis has no network dependency once the required models are installed, which is a condition, not a slogan. Read every row the same way: the score is about one axis, under the conditions stated in its basis.

The same six axes, drawn. On-device leads on privacy, offline reliability, cost efficiency, and customization; cloud leads on knowledge breadth, and on older hardware, on latency. Scores come from the rubric above and are our judgment, not benchmark results.

Privacy: on-device removes the copy

The privacy difference that matters is whether a copy of the conversation exists outside your devices at all. Encryption in transit is table stakes; the server copy is the architectural fact. On-device inference never creates the operator-side copy that a cloud service retains and can produce on request. There is no retention window to reason about, no training pipeline to opt out of, and no provider breach that exposes the journal, because the journal was never on the provider's systems.

What on-device does not remove is the risk around the device itself. Lose the phone with no backup and you lose the data. A backup you configure moves a copy somewhere else, and that somewhere joins your threat model. Anyone with your passcode can read what you wrote. On-device changes who holds your data. It does not make the device disappear.

Cloud privacy is real but conditional. Every major provider gives you controls for training, retention, and memory, and the controls do different things, which is why we read every policy and put the quotes in one dated table. What that work showed: cloud privacy is a policy question with settings. On-device privacy is an architecture question with no server copy to configure. Policies change. Architectures decide where the copy lives.

Scores: on-device 95, cloud 45. Basis: structural.

Offline reliability: one side keeps working

The network is not part of the on-device architecture. After the required models are installed, core functionality works in airplane mode, in a basement, on a mountain, and during the kind of outage that turns a cloud assistant into a retry button. No product team ships this deliberately. It falls out of where the compute lives, and it is the one axis where the comparison is not close.

A cloud model is a remote resource. Offline is not degraded service; it is no service. No provider can fix that without moving the model onto your device, because the capability is not local, and caching a request is not the same as running the model.

Two conditions keep the axis grounded. First, the models have to be installed before the flight, so the accurate claim is offline after setup, not offline out of the box. Second, offline does not mean identical. A device under memory pressure or thermal load can fall back to a smaller model, and the fallback changes quality. The axis still belongs to on-device, because the alternative on the same flight is nothing.

Scores: on-device 100, cloud 25. Basis: structural.

Latency: the network round trip is a stage you pay first

Latency is not one number. A reply moves through stages: speech recognition, embedding the query, searching memory, reranking results, generating the first token, finishing the answer, and optionally synthesizing speech. The stages have different owners in the two architectures, and the totals only make sense stage by stage.

On-device, every stage runs on the same silicon. Retrieval, meaning embedding, search, and reranking, costs a fraction of generation. The expensive stage is the first generated token, because the model has to read its context and start producing. Nothing waits on a radio, so on a supported device in a normal thermal state, the gap between finishing a sentence and seeing a response begin is dominated by the model.

Cloud adds a stage the local path does not have. The request travels out, waits its turn, and streams back. You pay for the round trip before anything visible happens, and the size of that cost is set by a network you do not control. After the first token, generation on a data center accelerator is very fast, and for long answers it leaves a phone model behind.

What a person feels is the gap between finishing a thought and seeing the reply begin. For short reflective turns, that gap is mostly network time, which is why on-device usually wins the comparison users actually make. For long generations, total time dominates and cloud usually wins. On entry-level hardware, the local model is smaller and slower, and the cloud lead widens. Thermal state is the last condition: a benchmark on a cool phone captures the best number the device will ever produce, which is exactly why it misleads. Twenty minutes of voice use looks different.

Per-stage latency in milliseconds for a short reflective query, with a mobile network round trip included in the cloud numbers. The first token and the round trip decide how the exchange feels; the retrieval stages are almost free by comparison.

Scores: on-device 78, cloud 62. Basis: stage-by-stage reasoning; the cloud lead grows on entry devices and long generations.

Knowledge breadth: cloud wins, and the gap is real

Models that fit in a phone are compressed, and knowledge is the first thing compression eats. Specific facts, names, dates, and niche domains degrade before reasoning does. The failure mode matters more than the gap itself: a compressed model often answers wrong at full confidence, and a wrong fact delivered without hesitation is a user problem you have to design around.

Retrieval is the counter-move, and it is specific about what it fixes. For questions about your own life, the knowledge that matters is your history, and searching that history locally narrows the practical gap to the point where a small model feels personal. That is the work behind the memory loop in Merrin, and it is why retrieval quality matters more than parameter count for this kind of use.

For questions about the world, retrieval from your history does nothing. If the workload is research, current events, or long-form reasoning across many sources, cloud wins that job today, and on-device should not pretend otherwise. A small model plus good retrieval beats a small model alone. It does not beat a frontier model at being a library.

Scores: on-device 55, cloud 95. Basis: judgment.

Cost efficiency: a bill versus a budget you already own

The two architectures have different cost shapes, and the shapes cross. Cloud charges per token, per seat, or per call. Light use is nearly free to start and stays cheap. Heavy use scales with how much people talk to the system, and at that point the bill becomes a permanent tax on the product. Rate limits, price changes, and enterprise minimums are vendor decisions that show up in your margin on the vendor's schedule.

On-device has no per-message bill. Its costs are fixed and mostly invisible: device capability, which the user may already own; storage for models and the local index; battery drain; and the engineering time to build a runtime, tune it, and keep it working across a device matrix. The last one is the expensive one. This article exists partly because that bill was paid here, and it was not small.

Cloud is cheaper until usage is heavy, and local is cheaper after that, provided the hardware in people's hands can run the model at all. The break-even point is specific to your product, which is why this axis carries a usage-dependent score instead of a single verdict.

Scores: on-device 90, cloud 70. Basis: usage-dependent; cloud wins at light use, local at sustained heavy use.

Customization: two different kinds of freedom

On-device gives you control of the stack. You choose the model, pin the version, tune the retrieval policy, and decide where data lives. No vendor can change the terms of software that runs on hardware you control. The price is maintenance: upgrades, tuning, and the support matrix are yours, and every new device tier is yours to test.

Cloud gives you a managed stack. New capabilities arrive without operations work, which is a real advantage for a small team. The catch is direction. The catalog is the vendor's, deprecations arrive on the vendor's schedule, and model behavior can shift under your product when a version retires. Teams that have been through an unplanned model retirement know what that costs.

Customization is freedom plus maintenance, and the deciding question is which one your team wants to own.

Scores: on-device 82, cloud 60. Basis: judgment.

Where cloud still wins

The subtitle promised this section, so here it is, without hedging.

  • Anything where the world is the corpus. Research, current events, unfamiliar domains, and general knowledge questions are cloud territory. A phone model is not a better library, and no amount of local processing changes that.
  • Long generations. On a good connection, cloud generation speed leaves phone hardware behind. If the product's center of gravity is long-form writing or analysis, cloud is the better engine.
  • Old and low-end hardware. If the device cannot hold the model, there is no comparison to make. Cloud runs on a ten-year-old laptop; on-device does not.
  • Zero operations. No model downloads to manage, no memory budgets to tune, no thermal engineering, no device matrix. For a small team, that is often the deciding factor, and pretending otherwise would be dishonest.
  • Elastic scale. Thousands of simultaneous users do not require provisioning. Cloud scales with demand; a device matrix does not.

If your data is allowed to leave the device, cloud is often simply the better engineering choice, and the axis scores say so. We are not going to pretend a phone model out-knows a data center. The interesting question is what the data boundary allows, and that is where the last two sections point.

Where this comparison breaks down

One caution before you quote any of this: cloud and on-device are classes, not products. Scoring them is like scoring the idea of a database. A well-tuned small local model and a careless hosted pipeline can flip several rows of the table, and your mileage will differ from the scorecard in ways that are specific to your build.

Variance is also high on both sides. Latency depends on the network, on radio state, on server load you do not control, and on the local device's thermal state. One clean run in a cool room is not a measurement of what your users see in week three.

The scores are a snapshot dated October 2026, and the two sides move at different speeds. Small models improve with every release cycle, and so do cloud models, which means the crossing points move. Re-run the rubric when the ground shifts, not this article.

Disclosure, because it belongs here: we build a local-first product. We have a horse in this race, and the rubric is the check on that bias. Every score names its basis, and the sections above say plainly where cloud wins, so you can disagree with a specific row instead of dismissing the whole picture.

How to decide for your own stack

Start from the data boundary, not the benchmark. If some category of data may never leave the device, cloud is off the table for that path, and the rest of the scores do not matter. That single constraint decides more architectures than any performance comparison.

The device ceiling comes next. Take the worst device you must support, not the best one in the office, and ask whether the model fits and whether the experience holds up. If it does not, you either raise the floor of supported devices or route that path to cloud. Both are legitimate answers. Pretending the middle device behaves like the flagship is not.

Then the shape of usage. Short, frequent, personal turns favor on-device, which skips the network and costs nothing per message. Long, occasional, knowledge-heavy requests favor cloud, which wins on quality and throughput. A product can have both shapes, which is how hybrid designs happen: route by the data boundary, then by the workload.

The decision order we recommend: data boundary first, device ceiling second, workload shape third, measurement last.

The last step is the one most teams skip: measure on your own hardware, on your own network, with your own workload. Our scorecard is a map of the terrain. It is not your terrain. If you are building in this space, the engineering deep dive covers how the memory loop makes a small model feel personal, and the build story covers what the on-device decision cost in practice.

Frequently asked questions

Is on-device AI better than cloud AI?

Neither wins the whole comparison. On-device leads on privacy, offline reliability, cost at heavy usage, and control. Cloud leads on knowledge breadth, long-form generation, old hardware, and zero operations. Which set matters more depends on your data boundary, your worst supported device, and the shape of your usage.

Is on-device AI as smart as cloud AI?

No. Models that fit in a phone are compressed, and they are weaker on knowledge-heavy questions, with confident wrong answers as the main failure mode. Retrieval narrows the gap for questions about your own history. It does not close the gap for general knowledge, where cloud models remain far ahead.

Is on-device AI faster than cloud AI?

It depends on the stage and the length. On-device usually feels faster for short conversational turns, because it skips the network round trip before the first token. Cloud usually wins long generations, and it wins more clearly on entry-level hardware. Thermal throttling and network quality can move both numbers a lot.

Does on-device AI mean my data never leaves the device?

Core data can stay on the device, but some network calls remain part of normal operation, such as model downloads and purchases. Backups and exports are user choices that move copies elsewhere. On-device removes the operator-side copy; it does not remove the risk of device loss, so keep your own backup.

Can you combine on-device and cloud AI?

Yes, and many products do. The practical rule is to route by data boundary first: anything that may not leave the device never goes to cloud, and anything allowed to leave can be routed by workload. Whatever crosses the boundary re-enters the world of retention windows and policy settings.

When should a company choose cloud AI?

When world knowledge, long-form generation, elastic scale, or zero operations matter more than the data boundary allows them to matter. If the data is allowed to leave the device and the workload is knowledge-heavy, cloud is usually the better engineering choice. Regulated and privacy-sensitive teams often choose on-device despite the quality trade, because the boundary is the requirement.