Three faults were leaving restaurant days out of balance — one in the data, two in the arithmetic — and a fourth, found late, that is a missing-data problem wearing a balancing problem's clothes. This is what they were, what they cost, and what fixing them is worth, measured by running the real job over ninety days of real trading, twice: once with the fixes off and once with them on.
In one sentence: a day's sales summary should show the money taken and the money earned agreeing to the penny, and on roughly one trading day in eight it did not — because two clients were fighting over the same records, tips that had been refunded were still counted as income, and service charges customers paid were credited to nothing.
How the two figures above were produced. Both are the real nightly job, run over the same ninety days against the same restored database, writing real summaries each time — the first pass with the fixes switched off, the second with them on. Comparing a re-run against a re-run rather than against production's stored summaries is the stricter test: production's figures are in places months stale, and crediting the fixes with repairing ordinary staleness would flatter them. On that fairer footing the fixes are worth a net 909 days and $60,784.96, not the larger number a stale baseline would have shown. Both passes ran with the duplicate client records left active, which is how this will actually be deployed.
Most of what is left is not a balancing fault at all, and the section on the fourth problem explains why deliberately leaving it unbalanced is the right call.
For the business: ten restaurant locations were set up twice in the system, as two separate clients. Both were importing from Square. Because the two records competed for the same payments and refunds, a refund would belong to one client for twenty minutes, then the other — so a day's books could gain or lose a refund depending on nothing but timing. On 2026-07-23 one client's summary was missing a $71.94 refund entirely, and was out of balance by exactly that amount.
technical Sales orders scoped their identifier by client (square/order/<code>-<loc>-<id>), but refunds, card charges, payouts and cash-drawer shifts did not — they used the bare Square id. Those attributes are :db.unique/identity, so both clients' imports resolved to a single entity and the last writer won.
Reading ownership out of the database's own history, this had actually happened to 3,387 refunds, 4,069 payouts and 2,628 cash-drawer shifts. And it has involved 19 client pairs, of which only 10 are visible in today's configuration — nine more contended in the past and the configuration has since changed, so no point-in-time check would find them.
For the business: the same collision meant a single card payment could be attached to both clients' copies of an order. That is worse than untidy. The nightly import removes orders Square reports as voided, and removing an order also removes its payments — so cancelling one client's order could silently delete the other client's payment, leaving a day showing sales with no money against them.
technical :sales-order/charges is declared :db/isComponent true, so [:db/retractEntity <order>] cascades into the charges. In a 20,000-order sample of the affected clients, 11,469 charges had two parent orders. This is why remove-voided-orders was left switched off during testing.
For the business: two arithmetic faults, both of which overstated or understated a day.
technical get-tip summed tips by joining through :sales-order/charges, so a return-only order — which has no tender to join through — contributed nothing, while its reversal sat unread on :sales-order/tip. Nothing at all read :sales-order/service-charge.
This is why the duplicated restaurants looked so much worse than everyone else. Of the days still failing after the first three fixes, 155 of the 423 on shared-location records had no sales orders at all — the summary consisted of nothing but refunds and their fees, with no sales for them to reduce.
The obvious reading is that the refund simply settled on a closed day. It is the wrong one. Checking each of those days against the date its client first recorded any order shows 140 of 171 fall before that client had a single order in the system — for seven of the nine records affected, every single one does. These are not quiet days. They are periods where the sales were never imported at all.
Where the refunds came from. Reading the database's own ownership history settles it. A $35.35 refund dated 26 February belonged to NGDG that same day, and was taken over by NGDU on 12 August. Others flip between the two records several times a day across 12–15 August. NGDU's first order is 2 August; it holds 94 refunds dated before it existed as a trading record. It never made them — it inherited seven months of the other record's refunds, because the refund key carried no client and whichever import ran last took ownership. That is fault 1, seen from the other end.
Across the nine records, 660 refunds worth $15,237.02 sit on a record dated before that record's first order. Nothing is lost and nothing is double-counted — the money is real and the surviving record has its own copy — but it is filed against a set of books that has no sales to put it against.
Why it is deliberately left out of balance. The day can be closed in one line: book a return equal to the day's refunds whenever the client recorded no sales. It is safe by construction — no trading day could be touched — and on this data it closes 171 of the 542 remaining days and $5,795.18. It was built, measured, and then removed, because it is the wrong thing to do. An unbalanced day is the only visible signal that a restaurant's sales are not being imported. Making the arithmetic agree would remove the alarm and leave the fire.
technical get-returns sums :sales-order/returns over orders scanned for the date. With no orders the sum is nil and no Returns line is written, while get-refund-items still credits Card Refunds from the sales-refund records. The imbalance is the correct output for the input; the input is what is wrong. A test now pins this behaviour in place so it is not "fixed" by someone reading only the arithmetic.
Five changes. The first three stop two clients from sharing a record; the last two record money that was being collected but not booked. Each is small — the difficulty was knowing which line to change, not writing it. A sixth was written and then removed; it is described at the end because the reasoning matters more than the code did.
Every imported record has an identifier the importer uses to decide "have I seen this before?". Sales orders already included the client; refunds, card payments, payouts and cash-drawer shifts did not, which is precisely why two clients could land on one record.
;; before — the bare Square id, identical for both clients (str "square/refund/" (:id r)) ;; square/refund/NOkQOTIiJULWN6… ;; after (scoped-key "square/refund/" client location (:id r)) ;; square/refund/NGCD-CD-NOkQOTIiJULWN6… (defn scoped-key [prefix client location id] (str prefix (:client/code client) "-" (:square-location/client-location location) "-" id))
Applied at five places in the Square importer: order payments, refunds, payouts (twice — the record itself and the lookup that finds it) and cash-drawer shifts. ezCater orders already did this and needed no change.
This is the one that makes the change safe to deploy. The identifiers are unique keys, so the importer relies on "same id, same record". Rename them and the next import matches nothing — and would quietly create a second copy of every refund and payment in the system, leaving the originals orphaned. So the importer looks up the record explicitly, new name first, old name second, and writes to whichever it finds.
(defn existing-id [db attr prefix client location id]
(when id
(or (dc/entid db [attr (scoped-key prefix client location id)]) ;; new scheme
(dc/entid db [attr (str prefix id)])))) ;; legacy scheme
The result is pinned as the record's id on the way in, so the write lands on the existing row regardless of which name it currently carries. The proof this worked is a count that did not move. Every one of the 265,965 refunds, payouts and cash-drawer shifts in the database was re-named, and afterwards there were still exactly 50,986 refunds, 144,688 payouts and 69,291 cash-drawer shifts — the same three figures as at the restore point. Had the fallback lookup been missing, each of these would have doubled instead.
The fallback also has to refuse. Reading the old name is what stops duplicates; reading anyone's old name is what creates them. Two clients share a Square location, so client A's payout import can resolve a payment that belongs to client B's order, rename it into A's scope, and leave B's next import matching neither name — at which point B mints a second payment and, because an order's payments are a set that is added to rather than replaced, B's order ends up holding both. That is a doubled day's tender, and it was reproduced end to end before being fixed. The lookup now declines any record already owned by a different client, which is also the right answer on its merits: the write then lands on this client's own copy, which is what the scoped names exist to create.
This is transitional. Once no legacy names remain, the fallback and the refusal are deleted together and the guarantee stops depending on either.
Renaming stops new collisions but does not undo old ones: a payment already shared by two orders is still one row with two owners. The migration walks each order's payments and, where another order has already claimed one, makes that order its own copy with the same amounts and points the order at the copy.
;; for each order, for each of its payments: :keep → first order to claim it; rename in place :clone → copy type, total, tip, tax, date, processor, note, receipt link set the copy's client and location to this order's retract this order's link to the shared payment link it to the copy instead
Run over the whole database that was 16,236,839 renamed and 500,438 copied, and payments owned by two orders went from 11,469 in a 20,000-order sample to zero across every order of the last year. The record count rose by about 500,438 — the number of copies it reported making, which is the check that it created what it meant to and nothing else.
One subtlety worth recording, because it bit us: the Square id has to be recovered from the record's current owner rather than by trimming a fixed prefix. Client codes contain dashes — N-30003 — so a pattern cannot tell where the client name ends and the Square id begins. Getting this wrong scoped some records twice and doubled their tender.
Tips were summed by walking from the order to its payments. A refund-only order has no payment attached, so its negative tip was invisible. The fix adds those tips rather than replacing the calculation.
;; before :ledger-mapped/amount (tendered-tip c date) ;; after :ledger-mapped/amount (+ (tendered-tip c date) (untendered-tip c date)) ;; untendered-tip — tips on orders with no payment attached [?e :sales-order/tip ?tip] (not [?e :sales-order/charges])
Adding rather than replacing is deliberate. Where an order does have a payment, the payment is the correct source: real orders exist whose payment carries a tip the order does not — an auto-gratuity recorded as a service charge, or a wallet tip missing from the order totals. Reading the order instead would have dropped those. Three tests hold this in place: the refund case must change, and the tendered and ordinary cases must not.
Nothing read the service-charge field at all. A new line credits it, for Square orders only and for negative amounts as well as positive.
[?e :sales-order/service-charge ?service-charge]
(or-join [?e]
[?e :sales-order/vendor :vendor/ccp-square]
(and (not [?e :sales-order/vendor])
[?e :sales-order/external-id ?external-id]
[(clojure.string/starts-with? ?external-id "square/order/")]))
Why the vendor test has two branches. ezCater service charges are commission the platform deducts from the restaurant, not money the diner hands over, so crediting them would make a day worse rather than better — hence the Square-only condition. But whole eras of Square orders carry no vendor field at all, and a test on vendor alone would silently credit nothing. The second branch falls back to the order's own identifier.
Why negatives matter. A returned catering fee arrives as a negative service charge and is already deducted from the day's returns; dropping negatives would lose the reversal. The line sits behind a per-client switch, off by default, so it can be turned on a few restaurants at a time.
Worth recording, because the arithmetic case for it is good and someone will propose it
again. Where a day has refunds and no sales orders whatsoever, book a Returns
debit equal to that day's refunds:
(defn- refund-only-returns [c date]
(when-not (traded? c date)
(let [amount (refunded-total c date)]
(when-not (zero? amount) amount))))
It works. Measured over the same ninety days it closed 171 days and $5,795.18, knocked nothing out of balance, and altered no already-balanced day — the guard makes it incapable of touching a day that traded.
It was removed anyway. Those days are not quiet days; they are days whose sales were never imported, and closing them removes the only visible sign of that. What is left in the code is a comment saying so and a test asserting the day stays out of balance, so the next person to notice the arithmetic finds the reasoning before they find the fix.
| Change | Why |
|---|---|
| Log each day's imbalance and its suspect lines | an out-of-balance day was only visible by opening the screen; now it can be queried |
| Stop the dirty-summary scan at the client boundary | it read every later client's summaries too — 1,321 ms to 5.6 ms per client |
| Split the recompute driver into a per-client function | lets a backfill spread clients across threads instead of grinding one at a time |
| Install schema attributes before the tuples that compose them | the test suite could not build an empty database at all, so no test could run |
That last one is worth a sentence for engineers: transact-schema installed
schema.edn then cloud-migration-schema.edn, but a composite tuple in the first file is built
from an attribute in the second. Datomic will not create a tuple before its members exist, so
every test fixture died in setup. It is very likely why sales summaries had no tests before
this work.
The job was run over the same ninety days at each stage, writing real summaries every time, so these are measured outcomes rather than estimates. All 18,900 client-day summaries in the window are included, whether or not the restaurant traded that day.
| Stage | Days out of balance | Clean | Total variance |
|---|---|---|---|
| Today's calculation, ninety days re-run | 1,451 | 92.32% | $81,023.96 |
| + refunded tips | 1,199 | 93.66% | $78,522.85 |
| + service charges | 542 | 97.13% | $20,239.00 |
Deduplication is not a row in this table, and that is deliberate. Separating the shared records is a change to the data, not to the arithmetic, and it had already been carried out before either pass ran — so both the baseline and the result above are computed on repaired data, and neither is credited with it. Its effect is shown structurally instead, further down: payments owned by two clients went to zero and stayed there. The consequence for reading this table is that $66,589.75 is what the three arithmetic fixes are worth on their own, with the deduplication's contribution already banked in the starting figure rather than added to the improvement.
| Change | Unchanged | Into balance | Out of balance | Balanced days altered | Money moved |
|---|---|---|---|---|---|
| Refunded tips | 18,588 | 252 | 6 | 6 | $3,777.67 |
| Service charges | 18,204 | 663 | 0 | 0 | $58,923.85 |
| Both, end to end | 17,916 | 915 | 6 | 6 | $60,784.96 |
Six days were knocked out of balance, and they are worth understanding rather than hiding. Across all 18,900 client-days, 17,916 summaries came out byte-identical and 978 of the 984 that moved were already wrong. The six exceptions all have one shape: the Tip line drops by a round amount — $10, $15, $20, $30, $30, $50 — and the day breaks by exactly that. They are all on shared-location records (NGDG, NGEZ, NGDA, NGBK).
That is the tip fix working, not failing. Each is a tip that was handed back: the reversal sits on this record's order, but the refund that should offset it went to the record's twin. Before the fix the day balanced by accident, because the reversal was ignored. After it, the day correctly shows that half the transaction is filed elsewhere. The honest description is that the fix converts a hidden mis-attribution into a visible one — which is the same trade the fourth problem below makes deliberately.
Service charges are by far the larger of the two fixes, moving $58,923.85 against the tip fix's $3,777.67.
That claim is stronger than a balance check, and it is the one worth insisting on: a day can stay balanced while its individual lines move, which would still be a change to the books. Every line of every summary was compared — category, debit or credit side, amount to the cent, and account — not just the day's bottom line.
The two fixes account for every day they moved, exactly. Adding up the untendered-tip and service-charge amounts across all 984 changed days leaves a residue of 0.0000000002. Nothing else moved those days; there is no unexplained remainder hiding a third effect, and the six that broke are accounted for by the same arithmetic as the 915 that healed.
Both arithmetic fixes add exactly one credit line. Nothing else in a summary moves — no sales figure, no payment, no tax.
| Line | Before | After |
|---|---|---|
| Tip | 482.94 | 422.94 |
| Card Refunds | 60.00 | 60.00 |
| Total money taken | 10,094.81 | 10,094.81 |
| Total money earned | 10,154.81 | 10,094.81 |
| Out of balance by | −60.00 | 0.00 |
The day already carried a $60.00 card refund — the guest was given their money back, tip included — while the tip line still credited the full $482.94. The corrected figure matches the refund to the penny. The order behind it is square/order/NGLK-SM-OxSX9gpXJV394qqT8mnBGypUwKNZY: a tip of −60.00 on an order with no payment attached at all.
| Line | Before | After |
|---|---|---|
| Service Charges | not shown | 427.10 |
| Card Payments | 4,975.89 | 4,975.89 |
| Total money taken | 7,777.20 | 7,777.20 |
| Total money earned | 7,350.10 | 7,777.20 |
| Out of balance by | +427.10 | 0.00 |
| Client | Date | Line | Before | After | Day closed |
|---|---|---|---|---|---|
| NGPA | 2026-06-04 | Service Charges | not shown | 1,344.86 | +1,344.86 → 0 |
| NTPT | 2026-08-06 | Service Charges | not shown | 427.10 | +427.10 → 0 |
| N-30003 | 2026-05-27 | Service Charges | not shown | 405.83 | +405.83 → 0 |
| NGFL | 2026-05-19 | Tip | 238.46 | 70.42 | −168.04 → 0 |
| NGMI | 2026-07-09 | Tip | 230.01 | 80.01 | −150.00 → 0 |
| NGVA | 2026-07-03 | Tip | 152.66 | 40.12 | −112.54 → 0 |
In every case the correction equals the imbalance exactly, which is what you would expect if the fix is recording something real that was recorded nowhere. On NGNP 2026-06-25 both fixes land on one day and pull opposite ways — $301.40 credited, $1.80 removed, $299.60 closed — a useful check that they are independent.
Every step below was performed against a restored copy of the production database. Production itself was never touched.
| Step | Result |
|---|---|
| Walk every order in the database, newest month first | 19,040,785 orders · 38 minutes |
| Give every order its own payment record | 16,236,839 re-keyed · 500,438 copied |
| Payments owned by two orders | 0 across every order of the last year — 5,158,470 |
| Client-scope refunds, payouts and cash-drawer shifts | counts unchanged · 0 collisions |
| Live Square import afterwards | 0 orders with duplicated payment · 0 shared payments |
| Ownership changes after the change | 0 refunds · 0 payouts · 0 shifts |
The count checks are the ones that matter. If re-keying had gone wrong it would have created a second copy of every record rather than updating the existing one, and the totals would have doubled. They did not move. The payment-copy step is the exception and is meant to add records — it added 500,438, matching the number of copies it reported making. (Close, not exact: the counter increments while the transaction is being assembled, so two copies that resolve onto one entity are counted twice. It is a good check, not a proof.)
The measurement above was taken with the duplicate client records deactivated, and that is not how this will be deployed. Leaving both records live is the intended configuration — the re-key is what separates them — but it means the deactivation that made this measurement clean will not be there. The gap that opens is narrow and specific: while any record still carries a legacy key, a second client can resolve onto it. That is why the deployment runs the migration with imports paused, and why existing-id now refuses to resolve a record belonging to another client.
Everything above was rebuilt from a fresh restore of the production backup, three times over, each time from the backup point itself rather than from a database an earlier run had touched: restore, re-key and split across all nineteen million orders, then two full ninety-day recomputes. The runs used different preparation — one deactivated the duplicate records and ran a live Square import, this one leaves the configuration exactly as production has it — so their headline figures differ, and that difference is itself the most useful measurement in this report. What did not move is the part that should not: for the 190 clients that do not share a Square location, the residue is 119 days and $1,730.61 in both runs, with the same five restaurants accounting for it. The arithmetic fixes behave identically no matter what is done to the duplicates.
The first attempt at copying shared payments derived each payment's Square identifier by stripping a fixed prefix. That is right the first time a payment is seen, but once it has been re-keyed to one client, a second order meeting it later read the already-scoped key as the identifier and scoped it twice — NGCD-CD-NGCC-CC-<id>. The importer then created a fresh payment, doubling the tender on five clients by $3,000–$7,000 each. It was caught because the totals were absurd, not because the code looked wrong. The fix recovers the scope from the record itself; client codes contain dashes, so it cannot be done by pattern. A test now runs the step one order at a time, which is the arrangement that exposes it.
542 client-days out of 18,900, totalling $20,239.00. Almost all of it sits on one group of restaurants, and that is the finding:
| Where the remainder sits | Days | Variance | Share |
|---|---|---|---|
| The twenty records that share a Square location | 423 | $18,508.39 | 91% |
| Every other client — 190 of the 210 | 119 | $1,730.61 | 9% |
For every restaurant that is not one of the ten duplicated pairs, this reproduces to the penny. 119 days and $1,730.61, with the same five clients accounting for almost all of it — NG4S at $1,066.61, NGMV at $259.38, NGEB at $199.09, NGPS at $172.82, N-30012 at $30.31, and 91 further days totalling $2.40 of till rounding. Those are the numbers an earlier run produced on a differently prepared database, which is a stronger check on the two arithmetic fixes than any single measurement: they behave identically whatever is done to the duplicates.
The $18,508.39 on the twenty shared-location records is the cost of leaving both records live. Each restaurant now keeps two sets of books, and the history behind them was never split: refunds claimed by whichever record imported them first, orders that went to the other, tips reversed on one side and refunded on the other. Re-keying makes that attribution stable — it stops moving — but it does not make it right.
Of that remainder, 155 days and $4,567.53 are days those records had no sales imported at all, which is the fourth problem above and is deliberately left visible. The other 268 days are real trading days on which the two records disagree about who owns what.
An earlier measurement of the same window, taken with one record of each pair deactivated and a live Square import run afterwards, left 279 days and $7,790.54 instead of 542 and $20,239.00. Most of that difference is the twenty shared records: with the duplicates retired their share fell from $18,508.39 to $6,059.93.
That is not an argument that the configuration is wrong — two live records is a deliberate choice, and the re-key is what makes it safe. It is a number to weigh: leaving both active costs roughly $12,000 of unexplained variance across 260 client-days per ninety days, carried on ten restaurants, until the historical attribution behind them is redistributed. The two figures are not perfectly isolated — that earlier run also included a live import, which backfilled data this one does not have — so treat it as the right order of magnitude rather than an exact price.
The importer understands both the old and new record names, so the change can be deployed before the renaming finishes. That tolerance is a bridge, not a destination — while any record still carries an unscoped name, two clients can land on it and the guarantee rests on a convention rather than on the data. So the renaming was run to completion and measured.
| Record type | Total | Client-scoped | Still to rename | Cannot be scoped |
|---|---|---|---|---|
| Card payments | 17,045,933 | 17,045,933 | 0 | 0 |
| Refunds | 50,986 | 50,986 | 0 | 0 |
| Payouts | 144,688 | 144,652 | 0 | 36 |
| Cash-drawer shifts | 69,291 | 69,291 | 0 | 0 |
Every record in the database now carries its owner's name, and the migration proposes no further changes: asked what is left to do, it answers zero on all four record types. The 36 payouts are ones with no client or location recorded anywhere, on the record itself or on anything referring to it, so there is nothing to name them after.
Renaming had to be driven from orders, because a payment's rightful owner is whichever order refers to it — so completing it meant walking all 19,040,785 orders, not just the clients that look shared today. Nine client pairs contended in the past without sharing a location now, and a migration scoped to the current configuration would have missed every one of them.
A caution about how completeness is counted. 282,649 card payments carry no client attribute of their own — they are stubs the payout path creates, never referenced by an order. A gate that checks the attribute reports these as "no owner" and looks like a gap. They are not: their names are scoped, recovered from the deposit that holds them. The figures above are counted the harder way, by asking the migration what it would still change, which resolves each record's owner through whatever refers to it. Reading the attribute alone would have understated completeness by a quarter of a million records — and an early draft of this report did exactly that.
| Shared payments after the migration | Count | Meaning |
|---|---|---|
| Owned by more than one order | 0 | whether the orders belong to different clients or the same one |
| checked across | 400,000 orders | spread through the whole database |
The problem this work exists to solve is gone: no payment answers to two orders, so the component relationship means what it says and deleting an order can no longer take another order's money with it.
Where Square splits one tender across two of a single client's own orders, the
payment stays shared — at any batch size. Both orders compute the same name, so there
is no second name for a copy to take, and a copy would double that client's takings for the
day. The mechanism is worth stating precisely, because it is not obvious from reading: once
the first order re-keys the payment it also writes the owner attributes in the same
transaction, so a later order recovers the bare Square id from those, computes the name the
payment already carries, and the guard (not= old new-key) drops the
row before any copy decision is reached. Verified by running the migration at a batch size of
one, which forces the two orders into separate batches: no copy is made.
An earlier draft of this report claimed the opposite — that such pairs would be copied once the batches split them — and flagged it as unmeasured risk to check before production. That was wrong, and it is recorded here rather than quietly deleted because it did real damage: an independent reviewer cited this paragraph as evidence and raised a defect that does not exist. A test now pins the behaviour at batch size one.
The guard on remove-voided-orders is still worth having regardless. It is
cheap, and it makes the safety a property of the deletion rather than of the migration having
been run first.
After the complete pass, asking the migration what it would change next returns nothing — 17,045,933 payments examined, none to rename, none unscopable. A record that already carries the right name is left untouched, so the migration can be stopped, resumed, or repeated without consequence.
Its speed is worth a note for whoever schedules it: the whole nineteen million orders were walked in about thirty-eight minutes, month by month from the current month backwards so that stopping early leaves the recent end done. An earlier attempt appeared to be transactor-bound and was projected at two days, which is why a previous run narrowed it to the analysis window. That diagnosis was wrong. The bottleneck was garbage collection in the process driving the migration — freeing held memory took an unrelated recompute from 17 client-days a minute to 4,515. Nothing about the database or the transactor needed to change.
What follows from the gate reading zero. unscoped-report counts
these figures on demand. Now that unscoped is zero across the board, the importer's
understanding of the old name form can be removed — at which point two clients sharing a
location becomes structurally incapable of producing a shared record, rather than prevented by
a convention that a future import could quietly break. That removal is the one remaining step
of this piece of work.
| Item | Who decides | Why it matters |
|---|---|---|
| Which client record survives at each shared location | the business | the newer record generally has no history before the split, so keeping it loses years of the location's books |
| Which revenue account service charges post to | accounting | currently 49000 Service Income, chosen so the work could be measured; it affects reporting, never whether a day balances |
| 659 refunds on records that have no sales for them | the business, then engineering | the top open item. $15,225.24 dated before the holding record's own first order. Either the missing sales get imported, or the refunds move to the record that has them — but the books cannot close until one of the two happens |
| Whether to correct records the wrong client already owns | the business | the fix stops future mix-ups; it does not retrospectively move records claimed while the configuration was shared |
remove-voided-orders | engineering | safe once no payment has two parent orders; worth guarding regardless so it detaches rather than deletes |
The production backup had not written a restore point since 2025-03-10 — roughly seventeen months — even though data files were still uploading daily. A backup you cannot restore from is not a backup. A fresh one was taken on 2026-08-14 and is what this work used.
The database server was sized for a toy dataset: a 2 GB cache against 27 GB of data. Worth checking what production is set to.
Slowness here was misdiagnosed twice, in the same direction. Both a recompute crawling at 17 client-days a minute and a migration projected to take two days turned out to be garbage collection in the client process, not the database or the transactor. Freeing held memory took the recompute to 4,515 client-days a minute — a factor of 265 — and the migration finished in well under an hour. The lesson generalises: before concluding the transactor is the bottleneck, look at the heap of whatever is driving it.
;; the restored database, untouched production as of 2026-08-14 22:52 (def conn (d/connect "datomic:dev://localhost:4337/integreat-prod-restore")) ;; the two orders behind the worked examples (d/pull (d/db conn) '[*] [:sales-order/external-id "square/order/NGLK-SM-OxSX9gpXJV394qqT8mnBGypUwKNZY"]) (d/pull (d/db conn) '[*] [:sales-order/external-id "square/order/NTPT-PT-KrMZzcon1cpQEJUyetErkBIpcdEZY"]) ;; the gate: no payment may have two parent orders (rk/charges-with-multiple-parents (d/db conn) orders) ;; => 0 ;; ownership history — which records ever changed client (->> (d/datoms (d/history (d/db conn)) :aevt :sales-refund/client) (filter :added) (reduce (fn [m d] (update m (:e d) (fnil conj #{}) (:v d))) {}) (filter (fn [[_ owners]] (> (count owners) 1))) count)
The comparison tool is committed as auto-ap.jobs.compare-sales-summaries. Unit tests: lein test auto-ap.jobs.sales-summaries-test auto-ap.square.core3-test auto-ap.jobs.rekey-square-external-ids-test.
as-of — a correction to how this was measured
The obvious way to audit a recompute is to read the database at a point before it and diff:
Datomic keeps every past value, so no snapshot is needed. That is what
compare-sales-summaries was built to do, and for summary amounts it does not work.
:ledger-mapped/amount, :ledger-mapped/ledger-side and
:ledger-mapped/account are all declared :db/noHistory true, so
superseded values are discarded rather than retained. A historical read of a summary that has
since been recomputed can return its lines with the categories intact and the amounts simply
absent — which reads as a legitimate all-zero summary, not as an error.
Every figure in this report is therefore taken from a live read of the database immediately after each pass, captured and stored outside it, and the before/after comparison is done between those two captures. No historical read is involved anywhere in the numbers above. The tool remains useful for categories and for which days changed; its docstring overstates what it can recover, and that is worth correcting before someone relies on it for amounts.