In short
- Problem
- Can a finance app understand your money without ever being told what it is — no account linking, no server, no upload?
- Constraint
- Zero bytes out. A phone rather than a datacentre, and a model that weighs about a gigabyte.
- Decided
- Rules first, model second. Stub the engine in phase 0 before it existed. Keep the native boundary small. Constrain generation, then validate it anyway. Converge PDF onto the CSV pipeline.
- Rejected
- Building the LLM layer first because it is the interesting part — an ordering that would have made every other phase wait on it.
- Hardest
- Running a language model inside a phone's memory budget, and giving the memory back afterwards.
- Outcome
- A working deterministic-first cascade at phase 4 of 5. Phase 5 is blocked on measurement from real use, deliberately. Unreleased, and no usage to report.
Question
Can an app understand your money without ever being told what it is?
India's financial life is UPI-heavy and SMS-native. A person's transaction history is scattered across bank SMS alerts — each bank with its own template and its own registered sender-ID conventions — UPI app notifications, PDF statements from banks and card issuers, and the occasional CSV export from a net-banking portal.
Every tool that stitches that together asks for the same thing in return: upload it all to our servers. And a transaction history is a startlingly complete portrait of a person — where they go, what they eat, who they pay, when they are struggling. It is a granularity most people would not hand to a stranger, and do, because the alternative is a spreadsheet.
So the question is not whether the categorisation can be done. It is whether it can be done without the data moving — whether the intelligence can come to the data instead of the data going to the intelligence.
Constraint
Zero bytes out. A phone, not a datacentre. And a model that weighs about a gigabyte.
The privacy rule is absolute rather than aspirational, and stating it that way removes every convenient option at once: no server-side parsing, no cloud model call, no analytics on transaction content, no "we only send anonymised data." Nothing leaves. That is the whole specification, and everything downstream is a consequence of it.
What remains has to run on a mid-range Android phone that is also running everything else its owner cares about. A resident model of roughly a gigabyte competing for that memory is a real risk rather than a theoretical one, and battery is a budget nobody approved.
The constraint is also the strategy. A privacy claim backed by architecture is verifiable in a way a privacy policy is not, and it is meaningfully harder to copy — matching it means solving on-device inference under mobile hardware limits, not calling a different API.
Decision
Eleven decisions. Several of them are about what to build before the interesting part.
Decision 01
Rules first. The model is the fallback, not the feature.
The parsing pipeline is a deterministic-first cascade. Layer one is regex and rule-based extraction handling the majority of cases — known SMS templates from major Indian banks, known statement table layouts — deterministically, fast, at zero inference cost and zero battery impact. Layer two is the on-device model, invoked only for the formats the rules cannot confidently parse, and for merchant-name normalisation where fuzzy semantic judgment genuinely helps.
This avoids a false choice. Handling every format perfectly with rules is unrealistic given how much SMS templates and statement layouts vary across dozens of Indian banks and card issuers. Routing everything through the model is unrealistic given what mobile inference costs in time and battery. The cascade takes the cheap path for the common case and reserves the expensive one for the tail.
Traded away: a single code path. Two parsing strategies means two sets of failure modes and a confidence threshold that has to be tuned rather than assumed.
Decision 02
Stub the engine in Phase 0, before it exists
An abstract engine interface, with a fake stub implementation, was defined in the very first phase — before any model work at all.
The effect is that the UI, the database layer, every parsing pipeline and the whole test suite could be built, run and tested without the real gigabyte model ever being downloaded or loaded. In a phased build where the inference subsystem lands late, defining that seam first is the difference between a project that progresses and one gated on its hardest problem.
Rejected: building the LLM layer first because it is the interesting part. That ordering would have made every other phase wait on the one with the most unknowns.
Decision 03
Keep the native surface small
The model is small and quantised, and it is downloaded on first run rather than bundled so the install stays a reasonable size. The native runtime it runs on is pinned to a specific release rather than tracking its main branch, so the native build stays reproducible and auditable.
Between that runtime and the app sits a hand-written bridge exposing a deliberately minimal API. Keeping that surface small and stable isolates the app from the runtime's internal churn, and makes the native boundary — the place where a mistake becomes a crash rather than an exception — small enough to reason about and test exhaustively.
Inference runs off the UI thread, so a slow or stalled generation never freezes the app. And the interface is deliberately runtime-agnostic: a different on-device runtime can be swapped in behind the same seam without the change rippling outward.
8 more decisions inside Continue the Arth case study Decisions 04–11 · complete pipeline · phase history · about 1,500 words Show all 8 remaining decisions ↓ Hide the full record ↑
Decision 04
Constrain the generation, then validate it anyway
Model output is forced into valid JSON matching the app's schema using a grammar at generation time — the model cannot emit malformed structure because the grammar will not let it.
And then a schema-validation pass runs afterwards regardless. That is belt and braces, and it is deliberate: a grammar constrains shape, not sense, and model output that is about to be written into someone's financial record is not data I am willing to trust on one guarantee.
Decision 05
Give the memory back when the model is idle
An idle-unload policy releases the model from memory when it is not actively being used. On a RAM-constrained Android device, a gigabyte held resident because the user might ask another question is a gigabyte taken from everything else they are doing — and the fastest way to have a finance app uninstalled is to make the phone feel slow.
Decision 06
Converge PDF onto the CSV pipeline instead of building a second one
PDF import handles password-protected statements and detects scanned, image-only PDFs that need different treatment from text-layer ones. Then it reconstructs the page layout into a grid — and feeds that grid into the same header-mapping and row-parsing components the CSV and Excel pipeline already used, unchanged.
Building a parallel PDF-specific parser would have been the obvious route and would have doubled the surface where the two paths could quietly disagree about the same statement. Instead a cross-format consistency test asserts that the PDF-derived and CSV-derived versions of the same transaction agree.
That test deliberately asserts on every field deduplication relies on, rather than the easier, more obvious surface fields of date and amount. The fields that matter are the ones a duplicate would hide behind.
Decision 07
Keep the date out of the duplicate check
Transactions arriving from different sources are deduplicated on what identifies the payment itself — and for transactions carrying a strong bank reference, the date is deliberately left out.
This was discovered, not designed. An SMS alert and a statement row describing the same real-world payment can disagree on transaction date versus posting date. A date-sensitive check would have silently created duplicate records for a meaningful fraction of transactions — silently being the operative word, because nothing would have errored and the totals would simply have been wrong.
It is the kind of rule that looks arbitrary in a diff and obvious in a bug report, which is exactly why it has tests around it.
Decision 08
Score the header row rather than assume it
Net-banking exports do not agree on where a table starts. So CSV and Excel import detects the header by scoring candidate rows instead of assuming a fixed position — and where the source data is ambiguous about whether a row is a credit or a debit, a direction-inferred flag surfaces in the import preview so the user can check the inference before it is committed to the database.
Showing an inference as an inference is the honest version of automation. The alternative silently books a refund as a purchase and lets the user find it in a monthly total.
Decision 09
Audit dependency health as a first-class criterion
Maintenance activity, licence and single-maintainer risk are reviewed before a package is adopted, not after it fails. Two real problems were caught and replaced mid-build because of it.
SMS access moved off a package flagged as a maintenance risk and onto a small native bridge of my own — which also made room to normalise the registered sender IDs Indian banks send from, so an SMS can be reliably mapped back to its bank.
Database encryption moved off a fork flagged as poorly maintained and single-maintainer, onto an actively maintained encrypted store. For a security-sensitive dependency in an app whose entire promise is custody, a single-maintainer fork is a structural risk, not a convenience trade.
Document handling was likewise chosen after auditing alternatives, on licensing and active maintenance. Encryption keys live in the platform's hardware-backed key storage rather than in application code or plain preferences.
Decision 10
Open every phase with a review block, not with code
Each phase begins with a design review and issue-fixing block before implementation starts, on the theory that design problems are cheapest at the moment they are noticed and most expensive once something is built on top of them.
It works. The two dependency problems above were caught this way. So were two open items in the current phase — including a diagnosability gap where the model load path has no async acknowledgment, so a silent failure from memory pressure or a corrupt file produces no clear signal. On its own that is a minor annoyance. Once user-facing features sit on top of the load path it becomes an opaque bug class, which is precisely why it was worth naming before that happened rather than after.
Decision 11
Let the phone decide the next phase
The next phase has to choose how the model gets invoked as a fallback: inline, blocking the interface while it runs; on-demand, triggered explicitly by the user; or background-batch, queued and processed opportunistically. Those are three genuinely different products.
That decision is blocked on measurement, deliberately. It will be made from prefill and decode tokens per second and peak resident memory, measured on representative hardware — not estimated. Which is why the current phase is not considered complete until those benchmarks exist.
Guessing here is cheap and wrong. If decode is fast enough, inline is the best experience and everything else is over-engineering. If it is not, inline is unusable and shipping it would be the kind of mistake that gets discovered by users.
The pipeline
Three sources in, one normalised record out.
| Source | Notable handling |
|---|---|
| Transaction SMS | Bank sender IDs normalised; the same extractors shared across the cascade |
| CSV / Excel | Header row found by scoring; an inferred credit or debit is shown for review before import |
| PDF statements | Password-protected and scanned statements detected; reuses the CSV pipeline, checked by a cross-format consistency test |
Every source converges on the same normalised record — date, amount, direction, bank, reference number, merchant — and then through the same deduplication. From there the app produces spending patterns, category breakdowns, recurring subscription detection, anomaly flags, cash-flow trends and budgeting suggestions, and answers natural-language questions over the user's own data. All of it on the device.
The asymmetry the product has to live with
Android SMS read access needs a runtime permission, and store policy around SMS-reading applications requires a narrowly scoped, user-facing justification — transaction alert parsing only, on-device, never transmitted. iOS has no equivalent capability at all.
That is not a gap to be closed; it is a structural difference. The iOS import experience is shaped toward PDF and CSV import and manual entry, and the product designs around that rather than papering over it with a feature that cannot exist.
On regulation: RBI guidance on financial data handling and the DPDP Act are the backdrop. On-device processing simplifies that posture considerably — there is no server-side storage or transmission to reason about. It does not eliminate the work. Permission scoping, on-device retention and user data-deletion controls all still have to be designed deliberately rather than assumed fine because they happen locally.
Phase history
Each phase opened with a review and closed with something testable.
| Phase | Built |
|---|---|
| 0 — Foundations | Encrypted schema, domain models, the engine interface with its stub, and SMS golden tests as the parsing correctness baseline |
| 1 — SMS parsing | The rules library, native SMS access, sender normalisation, the deterministic-first cascade |
| 2 — CSV / Excel | Statement import, header detection, cross-source deduplication, inferred-direction review |
| 3 — PDF | Password handling, scanned-PDF detection, layout reconstruction feeding the existing pipeline unchanged, cross-format consistency test |
| 4 — Native inference | The on-device runtime and its bridge, off-thread generation with constrained output, validation, idle unload |
| 5 — Invocation model | Design blocked on benchmarks by intent: inline, on-demand or background-batch, decided by measured throughput and memory |
Five SMS golden tests in Phase 0 are the smallest artifact here and among the most load-bearing. A golden test written before the parser means every subsequent change to the regex cascade has to justify itself against a fixed expectation rather than against whatever the parser happened to produce last week.
Outcome
- 0Bytes of financial data transmitted
- 3Import sources, one pipeline
- 4/5Build phases complete
- 108Tests passing, analyzer clean, at the native-inference checkpoint
The database is encrypted at rest as well as never transmitted, with its keys held in the platform's hardware-backed storage rather than in the app.
The parsing pipelines, the database layer, the UI and the test suite were all built and exercised against a stubbed engine before the real model was ever loaded — which is the clearest evidence that Decision 02 was worth making early.
There is no usage to report. Phase 5 is blocked on measurement from real use, which is the same fact from the other direction — the numbers the next phase turns on are numbers the app has not collected yet, and inventing them would defeat the phase.
A privacy policy is a promise. An architecture is a constraint. Only one of them survives a change of management.
flutter · dart · kotlin · on-device language model · native bridge · encrypted local storage