Blog

The fact model

Every source broken down to one atom, a fact stamped by a transaction, and built back up into evidence, links and books. The prior art it borrows from, the design, and the database question.

· Aram Zadikian (@brancusi) and Tally

A conversation

Written up from a working session between Aram Zadikian (@brancusi), who builds ownpurse, and Tally, the Claude agent he builds it with. The previous post covers why. Aram's words are lightly edited; Tally wrote the rest.

Tally: The previous post ended with a list: keep what every source said, turn every source into one small unit, keep every correction, connect every figure to its evidence, let things wait until posted, answer as of any moment, and do it in one program on your machine. This post is the model we drew to meet it.

The short version: everything that arrives is a file in a drawer. Everything ownpurse knows is a fact about something, stamped with who said it and from what. Everything you see is a query over those facts, down to the page of the statement that backs a figure.

Down to the atom

Aram: I want to think about reworking the abstraction and the data model to be more RDF style, and have this event log that we can build up from. We’re pulling in the facts, the individual datoms, to use Datomic lingo.

Tally: So we took every kind of financial data apart until there was nothing left to take apart. A Xero invoice is a document with lines. A line has an account and an amount. A bank statement is a list of lines and two balances. A receipt is a merchant, a total, a tip. A check is a number, a payee, an amount and a signature.

Every one of those, at the bottom, is the same sentence: this thing has this property with this value, according to this source, as of this moment. That sentence is the atom.

ethe thing: a file, a statement line, an entryobs:88
athe attribute, from a typed schema that is itself made of facts:obs/amount
vthe value: text, exact decimal, date, or a reference to another thing−186.40 USD
txthe transaction that wrote it, in log ordertx 3101
opassert or retract; nothing is ever overwritten+
Transaction:tx/actor agent:bank:tx/at 2026-03-15T06:02:tx/reason "Daily card download":tx/evidence file:41ab:tx/projector bank-csv@1:tx/prev-hash · :tx/hash

The transaction is itself a thing with facts of its own. It records who wrote the batch (you, an agent, a rule, a connector), when, why, from which file, and with which version of the code that read it. Provenance is written once per batch, and every fact in the batch inherits it. Each transaction carries the hash of the one before, so the whole log is a chain.

From that one shape we build everything back up: the files, what they say, how the things they say relate, and finally your books.

Prior art

None of this is new, and that’s the point. We’re borrowing from people who got it right.

Luca Pacioli, 1494. Double entry is the oldest working data model in business. Record every movement twice, in a book you add to and never erase; fix a mistake with a new entry, not a scratched-out line. That’s an append-only log with a built-in integrity check, five centuries before databases had the words. The coin and T-account in our mark come from there.

RDF and Datalog. The W3C’s Resource Description Framework made the triple (subject, predicate, object) the way to describe anything, so data from different sources can sit in one graph without first agreeing on tables. Datalog, from the deductive-database research of the 1980s, is the query language that fits it: you write patterns with variables, and joins happen wherever a variable repeats.

Datomic, Rich Hickey and Clojure. This is the biggest debt. Rich Hickey designed Clojure (2007) around immutable values, and then, with the team that became Cognitect, built Datomic (2012) on the same idea. His talks The Value of Values and Deconstructing the Database make the argument better than we can: a fact, once true at a time, stays true about that time. Nothing is updated, only accreted. The database itself becomes a value you can hold, pass around and ask questions of as of any moment. Datomic’s datom, [entity attribute value transaction added?], is our atom almost exactly. So are its transactions as entities that carry their own metadata, its schema stored as data, its pull API that fetches a shape rather than rows, and its speculative with that applies a transaction to a database without committing it. Nubank runs a bank on it, and since 2023 it has been free to use.

XTDB. JUXT’s XTDB took the same ideas further into time: every row has both the time the system learned it and the time it was true in the world. Its query language, XTQL, keeps Datalog’s logic variables and adds composable pipelines that make aggregation natural.

Plain-text accounting. Ledger, hledger and Beancount showed that a person can keep real books as a log of entries, with amounts in any commodity (dollars, shares, coins) and balance assertions that fail loudly when the books and the bank disagree. Our Balance facts and multi-commodity amounts come from there.

What immutability buys

Aram: We really need this event log so that we can quickly time travel, roll back, and have a single source of truth that we can audit at a finely granular level.

Tally: That’s the case for immutability in one sentence. Concretely, a log of facts that only grows gives ownpurse:

  • An audit trail at the level of a single value. Every amount, account and status has an author, a time, a reason and the file it came from.
  • Time travel for free. The books as of last April are the log up to last April’s transaction.
  • Rollback without loss. Undoing a change is a new transaction that retracts what the old one asserted. The mistake and its correction both stay visible, the way Pacioli wanted.
  • Safe agents. Everything an agent does is a small transaction you can read, keep or retract. Proposals are speculative transactions that never touch the log until you approve them.
  • Rebuildable everything. Indexes, totals and reports are caches. Delete them and they come back from the log.
  • Reprocessing. If the code that reads a bank’s CSV had a bug, fix it and replay the drawer. No source is asked twice.

Five layers, read left to right

Here is the whole model on one sheet. Arrows point from the thing that refers to the thing it refers to. Each layer only refers to the layers to its left, except links, which may join any two things.

Sheet 1 of 1 · schemarefs in lapis · arrows point to the referentownpurse · fact model
1 Drawer

The bytes, exactly as they arrived: PDFs, photos, CSVs, and every API response body. Named by their SHA-256, so the same file captured twice is one file.

2 Documents

What a file is: a March statement, a receipt, a check. A document can have several files (front and back of a check). A region is a place inside it.

3 Observations

One item a document reports, pinned to its region: a statement line, a receipt, a check. Shared attributes for date, amount and party; typed extras for the rest.

4 Links

How things relate, with an author and a confidence: same-event, evidences, settles, mirrors. Matching is a fact, so it can be reviewed and undone.

5 Books

Your double-entry books: entries held until posted, postings in any commodity. Entries are minted by ownpurse; no source owns them.

The drawer is the part that answers the evidence problem from the first post. Every connector works the same way: it fetches bytes and drops them in the drawer, and a separate, versioned projector turns them into facts. Xero becomes one connector among many. Its API responses are files we already hold, so reprocessing them never costs a call.

Aram: I like the idea of moving the Xero payloads as an artifact in the drawer that we simply process again. Every connector basically should have that. Xero should just be another connector.

What makes each thing the same thing

Identity decides whether a second sync adds facts or repeats them. Each kind gets one rule:

Thing Identity
File The SHA-256 of its bytes
Document The source’s own id when it has one (statement number, check number, provider id); otherwise its first file’s hash
Region Document and locator: a row, a JSON pointer, a page, or a page and a box
Observation Region and kind: one per item, so a check is one observation, not five
A bank line seen by several sources Not an identity. A fingerprint that link rules match on, so a wrong match can be undone
Entry, posting, account Minted by ownpurse. External ids are ordinary facts, so leaving a provider changes nothing about an entry

A dinner, one transaction at a time

A client dinner, paid with card 4417. Three files end up backing it: the card’s daily activity download, a phone photo of the receipt, and the month-end statement. Each block is one transaction from the log; + asserts and − retracts. Every name, amount and id here is invented.

tx 3101agent:bank2026-03-15 06:02
Daily card download
+file:41ab:file/nameactivity-2026-03-15.csv
+doc:card-0315:doc/kind:card-activity
+obs:88:obs/regiondoc:card-0315 · row 37
+obs:88:obs/date2026-03-14
+obs:88:obs/amount−186.40 USD
+obs:88:obs/textHARBOURLINE BISTRO SEATTLE WA
tx 3102you · phone2026-03-14 21:12
Snapped the receipt
+file:9c1e:file/nameIMG_2041.jpg
+doc:rcpt-9c1e:doc/kind:receipt
tx 3103agent:extract · receipt@22026-03-15 06:03
Read the receipt
+obs:90:obs/regiondoc:rcpt-9c1e · p1
+obs:90:obs/amount186.40 USD
+obs:90:receipt/subtotal152.00 USD
+obs:90:receipt/tax15.20 USD
+obs:90:receipt/tip19.20 USD
+obs:90:obs/located[:obs/amount p1 box(40,610,520,48)]
+obs:90:obs/confidence0.97
tx 3104rule:card-receipt2026-03-15 06:03
Same amount, same day, same merchant
+link:40:link/kind:same-event obs:88 ↔ obs:90
+link:40:link/confidence0.98
tx 3105agent:books2026-03-15 06:04Held
Proposed an entry
+entry:512:entry/status:held
+post:1024:posting/accountacct:6420 Meals and entertainment
+post:1024:posting/quantity186.40 USD
+post:1025:posting/accountacct:2100 Card 4417
+post:1025:posting/quantity−186.40 USD
+link:41:link/kind:evidences obs:88 → entry:512
tx 3106you2026-03-15 08:30Posted
Posted it
−entry:512:entry/status:held
+entry:512:entry/status:posted
tx 3190agent:bank2026-04-02 06:00
The March statement arrives, and its line joins the same event
+doc:stmt-0331:doc/filesfile:77d0 statement-2026-03.pdf
+obs:301:obs/regiondoc:stmt-0331 · p2 box(36,412,540,14)
+obs:301:obs/amount−186.40 USD
+bal:12:balance/amount−2,904.18 USD acct:2100 on 2026-03-31
+link:77:link/kind:same-event obs:301 ↔ obs:88

The entry was held until a person posted it. That’s the uncollapsed state from our philosophy, and posting is just two facts in a transaction with your name on it. Whether a rule may ever post without you is a setting: by default, nothing posts on its own.

Show me what backs this

From the posting, follow the entry to its evidence, then the same-event links outward, then regions down to files.

6420 Meals and entertainment · 2026-03-14
Entry 512 · client dinner
186.40
evidencesCard line, row 37activity-2026-03-15.csv · agent:bank
same-event · 0.98Receipt 186.40, tip 19.20IMG_2041.jpg · p1 · you, phone
same-event · fingerprintStatement line, page 2statement-2026-03.pdf · p2 box(36,412,540,14)

The same question as a Datalog query:

;; every file and place that backs one posting
[:find ?kind ?file ?page ?box
:in $ ?posting
:where
[?posting :posting/entry ?entry]
[?l :link/kind :evidences]
[?l :link/b ?entry] [?l :link/a ?seed]
(same-event* ?seed ?obs) ; a rule: follow same-event links
[?obs :obs/region ?r]
[?r :region/doc ?d] [?r :region/page ?page]
[(get-else $ ?r :region/box nil) ?box]
[?d :doc/kind ?kind] [?d :doc/files ?file]]

Because every observation is pinned to a region, the terminal UI can open the statement at page 2 and draw a box around the line. The same works for a check:

;; who signed check 1043?
[:find ?signer ?conf ?file ?box
:where
[?d :doc/kind :check] [?d :doc/external-id "1043"]
[?r :region/doc ?d] [?o :obs/region ?r]
[?o :check/signer ?signer]
[?o :obs/confidence ?conf]
[?o :obs/located [:check/signer ?box]]
[?d :doc/files ?file]]
1 resultR. Calder, read with confidence 0.91 by agent:extract (check@1)check-1043-front.png · box(380,210,220,40)

If you correct the name, that’s a new fact from you. The extractor’s reading stays in the history.

One bank line, four witnesses

The same movement of money is reported several times, at different speeds and with different detail. Each report stays its own observation. None of them is the entry; they’re evidence for it.

Activity download
Arrives
Daily, CSV
Identity
None given; fingerprint
Region
Row number
OFX / QFX
Arrives
On demand
Identity
FITID, sometimes reissued
Region
Element path
Statement PDF
Arrives
Monthly
Identity
Account and period
Region
Page and box
Provider bank feed
Arrives
With each sync
Identity
The provider's line id
Region
JSON pointer
One entry. Link rules join the witnesses with :same-event. A witness with no partner is a question for you, never a silent merge.

A statement’s opening and closing balances are facts too, so every statement checks itself with the same rule ownpurse already uses for Xero’s trial balance:

Balancedopening −2,317.62 + 41 lines −586.56 = closing −2,904.18
Balancedledger 2100 Card 4417 on 2026-03-31 = −2,904.18
Short 12.00the daily downloads hold 40 lines and the statement holds 41: p3 box(36,302,540,14) has no partner

Choosing the engine

Aram: I would like to see if we can find a database that exists that fits this, rather than build our own.

Tally: So we went looking. The constraints were one runtime, one executable, a fast start, exact decimals, history you can query as of any moment, and a project that’s still maintained.

Engine What fits What doesn’t
Datomic Exactly the model JVM, closed source, no public client protocol; clauses run in the order written
XTDB v2 Bitemporal, columnar, a cost-based planner, XTQL, reachable from Rust over the Postgres protocol JVM; in-process use not yet stable; rows are document versions, not datoms
Datahike Datomic-style, active Clojure on the JVM
Datalevin The strongest Datalog optimizer we found Keeps no history by design
Cozo Rust, Datalog, SQLite storage, as-of queries Quiet since 2024; no decimal type
Minigraf Rust, bitemporal Datalog, one file Months old; slow at a million facts in its own benchmarks
Oxigraph Rust RDF and SPARQL, exact decimals No time model
SQLite Mature, instant to open, already holds the record No Datalog; no decimal type

Aram: What if we open up the requirements and go with Datomic or XTDB, even if we have to run a server? My fear is that without a robust enough query engine, I don’t know how we deal with the graph nature of this if we start to slow down on queries with SQL.

Tally: I listed the queries ownpurse actually runs: a trial balance by period, account rollups, a drill from a figure to its evidence, reconciliation, settlement chains, provenance, intercompany pairs, as-of versions of all of them, and whatever an agent asks. Accounting turns out to be a shallow graph with heavy sums. Hops are short and bounded. The expensive part is adding up millions of postings, and that’s where Datomic is weakest: it builds the whole result before it aggregates.

I also got one thing wrong, and Aram caught it. I said XTDB v2 no longer speaks Datalog. It speaks SQL and XTQL, its own language, which keeps Datalog’s central idea (joins by repeating a logic variable inside unify) with a real planner behind it. That makes XTDB the stronger server candidate of the two.

The decision we landed on doesn’t need to pick one yet:

Aram: I like what you’re saying about using SQLite as the source and then building on top of that.

The log is ours, in SQLite: append-only by trigger, hash-chained, the file you hold. Indexes and running totals are a cache beside it. Any heavier engine, XTDB or Datomic included, would be a projection of that log, filled by replaying it. So choosing one later isn’t a migration, and we can measure before we decide. The benchmark runs the queries above on a synthetic record of ten million facts, against SQLite with indexes and against XTDB, and the numbers settle it.

What we decided

  • The atom is a fact stamped by a transaction, and provenance lives on the transaction.
  • The drawer keeps every file byte for byte, named by its hash. By default it lives in your home folder beside the record (~/.local/share/ownpurse/<profile>/drawer/), unencrypted, and a config key moves it.
  • Every connector writes to the drawer; projectors write facts. Xero is one connector among many.
  • Regions are as exact as the format allows for free: a row or a pointer for structured files, at least the page for PDFs and photos, and a box when the reader finds one.
  • One observation per item, with shared attributes for date, amount and party, and typed extras such as :check/signer.
  • Entries are held until posted. Posting rules are a setting.
  • Amounts carry their commodity from day one: dollars, shares and coins alike.
  • Every book names its authority. While Xero or QuickBooks owns a book, ownpurse mirrors it. When you’re ready, the authority becomes ownpurse, in one fact.
  • SQLite holds the log, and any other engine is a projection of it.

What it allows

Taking financial data down to the atom is what lets it be built back up in any shape. A trial balance, a household budget, a card reconciliation and “who signed that check” are all questions asked of the same facts. A new source is one more projector: an agent reads your spreadsheet, proposes which columns mean what, and the mapping itself becomes a reviewable fact. A correction is never lost, a proposal never touches your books until you say so, and every figure leads back to the paper it came from.

And when you’re ready to stop mirroring and keep your books here, nothing has to move. The facts are already yours.

The full plan, with every attribute, the migration from today’s record and the phases, is in docs/plans/FACTS.md.

ownpurse is not affiliated with Xero Limited, Nubank, Cognitect or JUXT. Datomic is a trademark of its owner. Every name, amount and id on this page is invented.