Files
integreat/docs/2026-08-15-sales-summary-rollout-plan.md
Bryce 79a4b457b0 docs: bring the rollout plan's later steps up to date with the backfill
The headline and the new backfill step were current, but step 11 and
"what this will not fix" still carried pre-backfill figures. The backfill
did not just change the totals — it changed step 11's population.

Before it, nine records held refunds dated before their own first order:
660 refunds worth $15,237.02. Seven of those were shared-location twins
whose refunds only looked orphaned because their orders had never been
imported; replaying the window gave them their orders and the refunds
stopped predating them.

Two are left, and they are a different case — neither shares a Square
location, so there is no twin holding the other half:

  NG4S  first order 2026-05-29  79 refunds  $2,180.08  10 days
  NGPS  first order 2026-05-26  81 refunds  $2,167.60   7 days

Step 11 now recommends trying backfill-history on them first, with a
window reaching back before their first order, since that is exactly what
resolved the other seven.

"What this will not fix" re-measured: 122 days / $2,379.45, of which 106
are real trading days ($1,151.80) and 16 are refunds with no sales
imported ($1,227.65). Only 3 of the trading days are on shared-location
records, all already diagnosed. The other 103 days and $502.96 have been
identical in every run of this analysis — deactivated, live, and
backfilled — and are the floor this work reaches.

Also adds the backfill's ~5.9 hour runtime to the up-front table, and
notes that "two entities per Square object" holds automatically for new
imports but needs step 5 for existing history.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 19:04:50 -07:00

420 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Sales-summary balancing — rollout plan
Steps to execute, in order. Every step is either reversible or verifiable before the next one
begins. The one behaviour change that alters a client's books is behind a per-client feature flag
that is **off by default**, so merging and deploying this branch changes nothing on its own.
Measured on a restored copy of production (backup point `209608347`), 210 clients over
2026-05-10 → 2026-08-07, with the duplicate client records left active exactly as they will be in
production: **1,191 client-days out of balance / $70,276.50 → 122 days / $2,379.45**, of which only
32 are above ten cents. 1,069 days came into balance, none broke, and no already-balanced day had a
figure altered.
Of the $2,379.45 left, just **$648.84 across 3 days** is on the twenty shared-location records. The
other 119 days and $1,730.61 belong to ordinary clients and have not moved across any run of this
analysis.
Getting the shared records there needs step 5 — a historical backfill from Square. Without it they
carry 423 days and $18,508.39, because re-keying stops the two records fighting but does not give
each its own copy of the refunds.
---
## Before you start
| | |
|---|---|
| Flag introduced | `summary-service-charges` — off by default |
| Migration to run once | `auto-ap.jobs.rekey-square-external-ids/migrate-all!` |
| Expected migration runtime | ~38 minutes for 19M orders on a warm cache |
| Backfill runtime (step 5) | ~5.9 hours for 90 days across the 20 shared-location records — an overnight job |
| Nothing here touches | invoices, payments, the ledger, or any client without the flag set |
**Client configuration is left exactly as it is.** Ten Square locations are configured against two
client records each, and both stay active. The re-key is what resolves them: once every record
carries its owner in its key, each client's import resolves only its own records and the two
records keep independent, stable histories. No "which record survives" decision is needed, and
nothing is deactivated.
The consequence to be aware of: each Square payment, refund, payout and shift at a shared location
becomes **two entities, one per client record** — by design. That is the stable end state, not a
duplicate to clean up. If any report or export aggregates across client records, one restaurant's
takings would be counted twice at that layer. Nothing in this work changes that either way.
That holds automatically for everything imported *from now on*, because the keys carry the client.
It does **not** hold for history: refunds, payouts and shifts already in the database exist only
once, on whichever record imported them last, and re-keying freezes that rather than evening it out.
Step 5 is what brings the existing history into the same shape.
**The only window of risk is between deploying and finishing the migration**, while legacy keys
still exist for a client to resolve. Steps 26 exist to make that window effectively zero.
---
## Step 1 — Guard `remove-voided-orders`
Do this before the migration, not after. `:sales-order/charges` is `:db/isComponent true`, so
retracting an order cascades into its payments. Until step 4 finishes there are still payments with
two parent orders, and deleting one client's voided order can take the other client's payment with
it.
Either leave `remove-voided-orders` switched off until step 4 verifies clean, or change it to detach
a payment that has more than one parent rather than delete it. Detaching is worth doing regardless —
it makes the safety a property of the deletion rather than of the migration having been run first.
See `docs/2026-08-15-remove-voided-orders-risk.md`.
---
## Step 2 — Pause the Square importer
**This is what makes the deploy safe, and it is easy to skip.** Steps 2 through 6 should be one
maintenance action, not separate days' work.
While legacy keys exist, `square.core3/existing-id` falls back to them — and at a shared location
that is the one code path that can reach across client records. Running the migration with imports
paused means no client is resolving keys while the keys are being rewritten, so the window closes
entirely rather than merely narrowing.
The migration itself takes about **38 minutes** for all 19M orders, so the pause is short — and
if you need it shorter, see step 6: you can resume imports before it finishes.
---
## Step 3 — Deploy the code
Deploy the branch. The flag is absent from every client, so:
- tips are calculated exactly as they are today,
- no `Service Charges` line is written.
The only changes that take effect immediately are the safe ones: imbalance logging, the
dirty-summary scan bounded to one client (1,321 ms → 5.6 ms per client), the schema-ordering fix,
and the importer's new client-scoped keys.
**The importer reads both key schemes**, so the deploy does not depend on the migration having
finished. Two protections cover the interval before it does: imports are paused (step 2), and
`existing-id` refuses to resolve a record that already belongs to a different client. Do not remove
the legacy lookup yet — see step 10.
---
## Step 4 — Run the migration
Run it immediately after the deploy, while imports are still paused.
```clojure
(require '[auto-ap.jobs.rekey-square-external-ids :as rk])
;; read-only first — no two entities may want the same key. `plan` does NOT return a
;; :collisions key; you have to hand its :new-keys to `collisions` yourself.
(rk/collisions (:new-keys (rk/plan (d/db conn) :charge/external-id rk/charge-prefix)))
;; => [] (anything else: stop, do not migrate)
;; then the whole thing
(rk/migrate-all! 2000)
```
`migrate-all!` runs this same check itself, on every attribute including charges, and throws
rather than transacting if it finds one. Running it by hand first just means finding out before
the 38-minute walk rather than partway through it.
Runs in about thirty-eight minutes over 19M orders. It is **idempotent and resumable** — a record that
already carries the right name is skipped, so it can be stopped and re-run without consequence.
**It is also ordered so that stopping early is survivable.** Refunds, payouts and cash-drawer
shifts go first — a quarter of a million records, seconds of work — so an interruption cannot catch
them half done. The long part then walks orders **a month at a time, from the current month
backwards**, logging `::month-complete` as each finishes:
```
::month-complete :month "2026-08" :rekeyed 118203 :cloned 2244
::month-complete :month "2026-07" :rekeyed 241887 :cloned 4611
...
```
That ordering is the recovery plan. If it dies, everything from the last logged month forward is
fully scoped — and that recent window is what the importer actually reads — so **you can resume
imports against a partially migrated database** and finish the older tail later. Walking oldest
first would have spent the first several hours on 2019 data no import will touch, leaving exactly
the wrong end done.
If you do resume imports mid-migration, the ownership guard in `existing-id` is what keeps the
unmigrated tail safe: a client cannot resolve onto another client's legacy-keyed record.
If it appears to crawl, the cause is almost certainly garbage collection in the process driving it,
not the transactor. That misdiagnosis cost two days of projected runtime during this work. Free
retained memory in the REPL and re-measure before changing anything about the database.
**Verify.** Two checks, doing two different jobs — run both.
**(a) Completeness, across everything.** `plan` must report nothing left to do, for all four
attributes:
```clojure
(dissoc (rk/plan (d/db conn) :charge/external-id rk/charge-prefix) :new-keys)
;; => {:total 17045933 :to-migrate 0 :already-scoped 17045933 :unscopable 0}
```
Read `:to-migrate 0` **and** `:unscopable 0`. This is the authoritative signal, and it covers all
17M charges.
`unscoped-report` is useful colour but is not the gate: its `:no-owner` column never reaches zero
for charges, because ~283k payout stubs carry no `:charge/client` of their own and it classifies
by attribute rather than by resolving ownership. Judge completeness by `plan`.
**(b) The safety gate for the cascade** — no payment may answer to two orders, or re-enabling
`remove-voided-orders` in step 9 can delete a payment another order still needs. Check **every**
order in the last year, with no sampling:
```clojure
(let [db (d/db conn)
cs (map first (d/q '[:find ?c :where [?c :client/code _]] db))
year (java.util.Date. (- (.getTime (java.util.Date.)) (long (* 365 86400000))))]
(rk/charges-with-multiple-parents
db (map first (iol-ion.query/scan-sales-orders db cs year nil))))
;; => 0
```
On the restored copy that is 5,158,470 orders — 27% of the table — via the
`:sales-order/client+date` index. A year is chosen deliberately: `remove-voided-orders` only ever
deletes orders Square reports as voided, which are recent, so that is where the destructive risk
lives. Completeness across all of history is check (a)'s job, not this one.
> Do **not** sample this with `(take n (rk/all-order-ids db))`. `all-order-ids` streams `:aevt`,
> which is ascending entity id, so a `take` returns the *oldest* orders — on the restored copy the
> first 400,000 are all from 20192021, before any of the contention this gate looks for. It would
> report a confident zero having inspected none of the relevant data.
---
## Step 5 — Backfill the shared-location clients from Square
**Skip this and the ten duplicated restaurants stay badly out of balance.** It is the difference
between 122 client-days out of balance and 542.
Sales orders have always been keyed by client, so both records of a pair built their own order
history. Refunds, payouts and cash-drawer shifts were not, so only ONE record holds each of them.
Re-keying freezes that ownership; it does not even it out. The record left without them shows
returns from its own orders and no refunds against them — NGBK held 158,535 orders and five
refunds — and is out of balance by exactly what its twin is holding.
Rather than manufacture copies, ask Square again. Client-scoped keys mean each record now creates
its own copy of whatever it reads, so replaying the window makes the two histories converge:
```clojure
(require '[auto-ap.square.core3 :as sq])
(require '[clj-time.core :as t])
@(apply sq/backfill-history
(t/date-time 2026 5 10) (t/date-time 2026 8 9)
["NGBK" "NGBR" "NGCD" "NGCC" "NGVG" "NGVC" "NGEZ" "NGJS" "NGDG" "NGDU"
"NGDV" "NGDS" "NGWC" "NGWN" "NGHY" "NGHA" "NGDA" "NGDL" "NGCL" "NGCT"])
```
**Verify** — every pair should hold matching order and refund counts in the window:
```clojure
;; per pair, per side: window orders and window refunds. The two sides should agree.
```
Measured on the restored copy: all ten pairs matched afterwards, and the shared records went from
423 days and $18,508.39 out of balance to 3 days and $648.84.
**Budget an overnight run.** This took **5.9 hours** for ninety days across the twenty records.
Every Square call in the process shares one 25-requests-per-second throttle, refunds and shifts cost
one API call per record, and `backfill-history` imports three clients at a time — raise its
`s/buffer` if you need it faster. Neither the database nor the transactor is the limit; reads
measured at 32 µs.
It must run **after** the migration. Run before, and it imports against legacy keys and leaves more
to migrate.
---
## Step 6 — Resume the Square importer
Normally: once step 4's two checks read clean and step 5's backfill has finished. The maintenance
window ends here.
**If the migration did not finish**, you do not have to wait for it. Resume imports once the
`::month-complete` log covers the window your importer reads — the last 75 days for payouts and
cash-drawer shifts, and whatever range the order import is configured for. Then re-run
`migrate-all!` afterwards to walk the remaining older months; it will skip everything already done.
Run the step 4 checks again once it does finish.
The first cycle after resuming is the one to watch. Compare these against the same counts taken
immediately before the deploy — growth should be ordinary daily volume:
```clojure
(count (d/datoms (d/db conn) :aevt :sales-refund/external-id))
(count (d/datoms (d/db conn) :aevt :expected-deposit/external-id))
(count (d/datoms (d/db conn) :aevt :cash-drawer-shift/external-id))
(count (d/datoms (d/db conn) :aevt :charge/external-id))
```
A near-doubling of any of them means records are being created rather than matched — **stop and
roll back the deploy.** Charges are included deliberately: they are the one that doubles a client's
takings rather than merely duplicating a row.
---
## Step 7 — Recompute summaries, flags still off
```clojure
(require '[auto-ap.jobs.sales-summaries :as ss])
(ss/refresh-sales-summaries 90)
```
This is the pass that banks the deduplication. **Capture the result before going further** — you
will need it as the baseline for step 8, and it cannot be reconstructed afterwards:
```clojure
(require '[auto-ap.tools.compare-sales-summaries :as cmp]) ; test/dev classpath
(def before (cmp/summaries-in (d/db conn) start end))
(spit "before.edn" (pr-str before))
```
> **Do not use `d/as-of` to compare summary amounts.** `:ledger-mapped/amount`, `ledger-side` and
> `account` are `:db/noHistory`, so past values are discarded. A summary that has since been
> recomputed reads back through `as-of` with its amounts *absent*, which looks like a legitimate
> balanced day. Capture live, before and after, and diff the captures.
---
## Step 8 — Turn the flag on, a few restaurants at a time
Needs accounting sign-off first: `summary-service-charges` posts to **49000 Service Income**, chosen
so the work could be measured. It affects reporting, never whether a day balances.
```clojure
@(d/transact conn [{:db/id [:client/code "NGxx"]
:client/feature-flags ["summary-service-charges"]}])
(ss/refresh-sales-summaries 90)
```
Start with two or three restaurants, confirm, then widen.
**Verify** against the capture from step 7:
```clojure
(def after (cmp/summaries-in (d/db conn) start end))
(cmp/compare-window ...) ; both arguments live database values, never as-of
```
The two numbers that matter — both were zero across all 18,900 client-days in testing:
- `:balanced->unbalanced` must be **0**
- previously-balanced days whose lines changed must be **0**
If either is non-zero, retract the flag for the affected clients and re-run step 7. The flag is the
rollback: removing it restores today's behaviour exactly.
---
## Step 9 — Re-enable `remove-voided-orders`
Safe once step 4's gate reads zero. Keep the detach-rather-than-delete guard from step 1.
---
## Step 10 — Remove the legacy key lookup
**Schedule this; do not leave it open-ended.** Both client records at a shared location stay active
permanently, so the legacy fallback in `square.core3/existing-id` is the one code path that can ever
reach across them. Deleting it is what turns the guarantee from conventional into structural.
Once `plan` reports `:to-migrate 0` and has stayed there through several import cycles, drop the
legacy branch of `existing-id` — and with it `owned-by-other-client?`, which exists only to make
that branch safe while it lives. After this, two clients on one location are structurally incapable
of resolving onto each other's records, and no ordering discipline is required to keep it that way.
Until it is done, the protection is the guard plus the maintenance window, both of which depend on
people doing the right thing. That is the reason not to let this drift.
---
## Step 11 — Deal with the refunds that have no sales behind them
**The most important item in this document, and the only one that is not just execution.**
16 of the 122 remaining days are a record carrying refunds on a day it recorded no sales at all,
and all 16 fall before that client's first ever order. Two clients are affected, holding **160
refunds worth $4,347.68 dated before their own first order**:
| Client | First order | Refunds before it | Value | Days out of balance |
|---|---|---:|---:|---:|
| NG4S | 2026-05-29 | 79 | $2,180.08 | 10 |
| NGPS | 2026-05-26 | 81 | $2,167.60 | 7 |
**Step 5's backfill already resolved the other seven.** Before it, nine records were in this state
holding 660 refunds worth $15,237.02 — but seven of them were shared-location twins whose refunds
only looked orphaned because their orders had never been imported. Replaying the window gave them
their orders, and the refunds stopped predating them.
NG4S and NGPS are different: neither shares a Square location, so there is no twin holding the other
half. Their sales genuinely are not in the system for the period their refunds cover. The database's
own ownership history is the evidence to check — for the twins it showed refunds changing hands
between the two records; for these two there is no second record to have taken them from.
Two ways to close it, and the business has to pick:
1. **Import the missing sales.** Correct if these records are meant to have their own books. Try
`backfill-history` for them first, with a window reaching back before their first order — that is
exactly what fixed the seven, and it is one command.
2. **Move the refunds to the record that has the sales.** Correct only if the refunds were misfiled
onto a record that should not have books of its own.
Start with (1): it is cheap, reversible in the sense that it only adds what Square reports, and it
is already proven to work on this exact symptom.
```clojure
;; per client: refunds dated before that client's own first order
(let [first-order (->> (d/q '[:find [?d ...] :in $ ?c
:where [?o :sales-order/client ?c] [?o :sales-order/date ?d]]
(d/db conn) [:client/code "NG4S"])
(reduce (fn [a b] (if (.before a b) a b))))]
(->> (d/q '[:find [(pull ?r [:sales-refund/date :sales-refund/total]) ...] :in $ ?c
:where [?r :sales-refund/client ?c]]
(d/db conn) [:client/code "NG4S"])
(filter #(.before (:sales-refund/date %) first-order))
count))
```
**Until this is resolved those days stay out of balance, on purpose.** A summary change to close
them was written and measured — it works, closes 16 days and $1,227.65, and breaks nothing — and it
was removed, because an unbalanced day is the only visible signal that a restaurant's sales are not
being imported. A test asserts the day stays unbalanced so nobody closes it without reading this.
---
## What this will not fix
122 client-days over ninety days, $2,379.45, of which only 32 are above ten cents.
| | Days | Variance | |
|---|---:|---:|---|
| Real trading days with genuine discrepancies | 106 | $1,151.80 | see below |
| Refunds on a record with no sales imported | 16 | $1,227.65 | step 11 — deliberately visible |
Of the 106 trading days, only **3 are on shared-location records** — $648.84 in total, and all three
are already diagnosed: NGBK and NGBR at $299.42 each on 2026-08-06, where Square recorded $6,358.99
of tender against $6,059.57 of order totals (the gap itself, not a summary fault), and NGDA at
$50.00, an auto-gratuity booked as a service charge.
The other 103 days come to **$502.96 across 190 clients** — a few dollars here and there, mostly
till rounding, plus small undiagnosed clusters on NGMV ($259.38 over 5 days) and NGEB ($199.09 over
4 days, an ezCater fee-treatment question). Those two are worth a look but are not urgent.
That 103-day, $502.96 figure has been identical in every run of this analysis — with the duplicates
deactivated, with them live, and with them backfilled. It is the floor this work reaches.
---
## Two operational findings, unrelated to the summaries
- **The production backup had not written a restore point since 2025-03-10** — about seventeen
months — although data files were still uploading daily. Worth an alert on restore-point age.
- **The database server is sized for a much smaller dataset**: a 2 GB cache against 27 GB of data.
Worth checking what production is set to.