Skip to content

feat(fygaro): float runway alerts + stranded-credit retry sweep - #516

Merged
islandbitcoin merged 9 commits into
mainfrom
feat/fygaro-float-runway-retry
Oct 4, 2026
Merged

islandbitcoin merged 9 commits into
mainfrom
feat/fygaro-float-runway-retry

Conversation

@islandbitcoin

@islandbitcoin islandbitcoin commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Problem

On 2026-10-03 the bankowner treasury (the Fygaro auto-credit funding source) ran dry. A $70 card top-up was captured, the credit failed with insufficient-treasury-float, and the customer was charged with nothing delivered until ops hand-credited $66.02 and clicked Mark Completed.

Three gaps made that worse than it needed to be:

  1. The warning was ignorable. The float monitor fired a warning every day for seven weeks (balance under $2000 since the day after the last refill). It carried no runway figure, never escalated, and shared nothing with the per-event CRITICAL that only fires after a customer has been charged.
  2. The cadence was not ours. The check ran from the cron Job, whose schedule the chart owns. In prod that was daily at 02:00.
  3. A failed credit never retried. payment.ts deliberately leaves the row Fiat Received with no failure_reason to invite a provider retry, but Fygaro does not redeliver a 200. Every float failure was a manual credit + a manual status stamp.

What changed

Float monitor (float-monitor.ts)

  • Computes runway = balance / trailing-7-day net credited (ERPNext Completed rows, final_amount). Burn read only when balance < 2x floor.
  • Two tiers: warning below floorUsd; critical below criticalFloorUsd (default $500 = the auto-credit single-payment limit) or under criticalRunwayDays (default 3). Critical has its own dedup key and its own Redis marker so an open warning window cannot swallow it.
  • Alert detail/context carry balance, floor, daily burn, runway, and the fund URL (ERPNext System Accounts page).
  • ERPNext is an input to runway only. A failed or throwing burn read still fires the floor alert with runway "unknown".
  • Returns the reading so the loop can decide whether to sweep.

Stranded-credit sweep (stranded-credit-sweep.ts, new)

  • Candidates: ERPNext rows provider=Fygaro, Topup, Fiat Received, USD, attributed, no failure_reason, not email-attributed, within retry.lookbackDays, oldest first, capped at retry.maxPerSweep. Rows refused by a gate carry a failure_reason and are never touched.
  • Processed markers are read first, before any gate or coverage check: the ERPNext row (ops may have hand-completed it) and a durable Redis "credited" marker (credited-marker.ts, written the moment a send succeeds, 30-day floor TTL decoupled from lookbackDays). A marker-credited row that is still not Completed pages critical ("promote by hand") and is never re-sent.
  • The FULL webhook gate is re-run (evaluateCreditGate over live settings, account level, trailing-24h gross excluding this row, operator fee-discount whitelist) — including the daily cap. Why: payment.ts answers 500 with no stamp for the transient settings-unavailable / history-unavailable refusals, so a candidate may never have passed the cap at all; crediting it blind could put a customer over a compliance cap with no refusal row and no alert. Outcomes:
    • transient reasons (settings-unavailable, history-unavailable): no stamp, retried next tick;
    • deterministic refusals (daily-limit-exceeded, over auto-credit limit, under minimum, non-USD, non-positive net, level with no cap): failure_reason stamped exactly as the webhook stamps it, warning alert per payment, row left for manual (and dropped from the daily-allowance sum). A failed stamp pages critical.
  • Send is the same creditFygaroTopup (idempotent on fygaro:<transactionId>) with the gate's own fee figure.
  • Coverage: availableUsd is balance minus the critical floor (the loop passes the reserve-adjusted figure). A row whose net exceeds what is left is reported uncovered (warning, per payment, "available above reserve") and skipped so a big row never blocks a small one. A float failure from the credit itself stops the sweep and pages fygaro:float-exhausted.
  • Aging out is never silent. Candidates are windowed on last_seen_at, which a failed or skipped retry never touches, so a row that keeps failing drops out of the list on day lookback+1. Every sweep also counts the same-shape rows older than the window (countAgedOutUncreditedFygaroTopups) and pages critical once per dedup window with the count and oldest request_id ("credit by hand").
  • On success: Redis marker written before the ERPNext promotion; row promoted with the full fee split (no more "Pending" fees on hand-completed rows); checkout intent stamped credited when the payload carries one; ops-feed success tagged retry: stranded-credit-sweep; customer pushed once under the same fygaro-credit-push: claim the webhook uses. A failed promotion after a successful send pages critical and is not retried by re-sending.

Treasury loop (treasury-loop.ts, new; started from webhook-server/index.ts)

  • Float check + sweep every fygaro.float.checkIntervalMs (default 15 min) inside the long-running fygaro-webhook pod. Redis NX lease per tick so one replica runs. First run 30s after boot. Sweep only when balance >= critical floor, and the sweep is handed balance − criticalFloorUsd: the floor is a reserve for live card traffic, so a backlog can never drain the float through it.
  • The cron Job keeps calling the float check too; both paths dedup through the same Redis markers.

Config (schema.ts, schema.types.d.ts, base-config.yaml) — all defaulted, no override required to deploy:

knob default
fygaro.float.floorUsd 2000 (unchanged)
fygaro.float.criticalFloorUsd 500
fygaro.float.criticalRunwayDays 3
fygaro.float.checkIntervalMs 900000
fygaro.float.fundUrl https://erp.flashapp.me/app/system-accounts
fygaro.retry.enabled true
fygaro.retry.lookbackDays 7 (must stay below the 30-day credited-marker TTL; see schema comment)
fygaro.retry.maxPerSweep 20

ERPNext readers (ErpNext.ts): sumFygaroCompletedNetCentsSince, listUncreditedFygaroTopups, countAgedOutUncreditedFygaroTopups. Reads only, existing fields only.

Deploy notes

  • No frappe-flash-admin change required. Both readers use fields already on Bridge Transfer Request.
  • No chart change required for the defaults. Set fygaro.float.fundUrl in the TEST overlay if the alert should point at the TEST ERP.
  • The fygaro-webhook workload needs Redis reachable (it already does, for checkout intents). Lease and markers are fail-open: a Redis blip means duplicate alerts, never silence.
  • critical severity routes wherever alertBridge sends criticals today. Confirm that path pages a human; the whole point is that warning did not. New criticals from this PR: float critical, float exhausted during retry, credited-but-not-promoted, refusal-not-stamped, aged-out stranded rows.
  • Do not raise fygaro.retry.lookbackDays to 30 or above without first raising MIN_CREDITED_MARKER_TTL_DAYS in credited-marker.ts.
  • First deploy will sweep any Fiat Received / no-reason / attributed Fygaro rows from the last 7 days. As of 2026-10-03 there were none outstanding (verified against prod ERP), so expect a no-op.

Test plan

  • yarn tsc-check clean, eslint clean on touched files, prettier --check (3.6.2) clean, typos clean.
  • Full unit suite: 277 suites / 3487 passed (3 skipped).
  • New/extended specs: float-monitor.spec.ts (30: runway math, tiering, separate critical marker, ERP fail-open on error and on throw, defaults), stranded-credit-sweep.spec.ts (gates incl. ERPNext kill-switch, processed markers first, full gate re-run incl. daily cap with transient-vs-deterministic outcomes, fee math + whitelist, reserve coverage ($600 float / $500 floor / $280 row → uncovered), float decrement, stop-on-float-exhausted, per-payment failure alert, aged-out paging incl. zero-candidate and error paths, never throws), credited-marker.spec.ts (new: 30-day TTL floor independent of lookback, fail-open write, unknown-on-read-error), treasury-loop.spec.ts (tick gating on critical floor, balance − floor handed to the sweep, boot delay + interval, lease TTL, lease held skip, Redis fail-open, clamp, survives throw), ErpNext.spec.ts (+20 for the three readers), BridgeTransferRequestWriter.spec.ts (+6), webhook-server/index.spec.ts (+1 boot wiring).

Follow-ups (not in this PR)

  • Gate fygaroCheckoutCreate on float coverage so a card is never charged for a credit that cannot be delivered. Inert until mobile calls that mutation (ENG-443/444).
  • Auto-rebalance bankowner from funder/dealer via the existing transfer_between_system_wallets.
  • ENG-545 reconciliation remains the backstop for payments that never reach the webhook at all.

🤖 Generated with Claude Code

bobodread876 and others added 9 commits October 4, 2026 12:08
…credit retry

Adds fygaro.float.{criticalFloorUsd,criticalRunwayDays,checkIntervalMs,fundUrl}
and fygaro.retry.{enabled,lookbackDays,maxPerSweep}, two ERPNext readers
(trailing net credited = treasury burn; Fiat Received rows with no
failure_reason = stranded credits) with their writer wrappers, and dedup keys
for the new alert classes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The daily $2000 warning ran for seven weeks before the 2026-10-03 exhaustion
and nobody acted on it. The check now computes runway from the trailing
7-day net credited, escalates to CRITICAL (own dedup key + own Redis marker,
so an open warning window cannot swallow it) below the critical floor or
under the runway threshold, and puts balance/burn/runway/fund-URL in the
alert. ERPNext is an input to runway only: if that read fails or throws the
floor alert still fires with runway "unknown". Returns the reading so the
treasury loop can decide whether to sweep.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rkload

A credit that fails on an empty treasury leaves the row Fiat Received with no
failure_reason and never retries (Fygaro does not redeliver a 200), so every
one became a manual credit + Mark Completed. The sweep re-runs the SAME
idempotent credit path for those rows once the float is back: same fee engine
and discount whitelist, status re-read right before the send so a hand-
completed row is skipped, promotion with the full fee split, per-sweep cap,
oldest first, rows refused by a gate never touched.

The loop runs float check + sweep every fygaro.float.checkIntervalMs inside
the fygaro-webhook pod under a Redis NX lease (replica-safe). The cron Job
keeps calling the float check, but its schedule is owned by the chart (daily
in prod), so the cadence that matters lives here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sweep:
- Durable credited marker (Redis `fygaro:sweep-credited:<tx>`, TTL
  lookbackDays+1) written after a successful send, before promotion or any
  announcement, by both the sweep and the webhook. The sweep checks it
  alongside the ERPNext row before every send, so a row whose promotion
  failed is never re-paid once the 24h send-idempotency cache expires.
  Promotion failure now pages CRITICAL ("money moved, promote by hand").
- Re-run the real credit gate (evaluateCreditGate over the live settings,
  account level, trailing-24h gross excluding the row, fee-discount
  whitelist). Deterministic refusals are stamped with
  markFygaroTopupNotCredited and raise the fygaroNotCredited warning;
  transient ERPNext reads defer to the next tick. The send uses gate.fees.
- Stamp the checkout intent `credited` when the stored payload carries an
  intent id, so fygaroTopupStatus agrees with the push.
- Add `leftForManual` to the summary for skipped/refused rows.

Float monitor:
- Read trailing burn on every tick; drop the nearFloor gate that made the
  runway tier dead exactly when burn was high.

Tests cover each: tick-2 no-resend after a failed promotion, marker order
before promote, daily-cap refusal stamped, intent stamped vs skipped,
leftForManual counts, and critical-on-short-runway at $5000.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…-but-not-promoted; marker-guard webhook re-deliveries

Review fixes for the stranded-credit sweep and the payment webhook:

- Sweep: read the ERPNext Completed row and the Redis credited marker
  directly after the transaction/account validation, before the credit
  gate and the float-coverage check. A row whose money already moved
  (hand-completed, or credited with a failed promotion) is no longer
  re-judged with today's inputs, so a lowered cap / raised minimum /
  downgraded level can no longer stamp daily-limit-exceeded or
  under-minimum on a paid row, drop it out of the allowance sum, or
  re-raise "uncovered" every tick.
- Sweep: the marker-credited branch now pages critical under
  erpnextFygaroAudit(transactionId) ("promote the row by hand") instead
  of only logging. Fygaro never retries a 200 and the sweep refuses to
  re-send, so this was the only path that could page a human.
- Webhook: the duplicate guard reads the credited marker after the
  Completed-row check. A re-delivery of a promotion-failed payment past
  the 24h send cache (ops resending from the Fygaro dashboard) now acks
  already_processed, stamps the intent credited with the marker's net,
  and pages critical instead of running creditFygaroTopup for real.
- Webhook: the promotion-failure comment/title no longer claims a
  provider retry will re-promote the row.

Specs: flipped the "completion AFTER balance check" assertion, added
marker-credited + gate-would-refuse -> no stamp, marker-credited -> critical,
Completed/marker rows never reported uncovered, and webhook marker-guard
(credited -> no send; unknown -> falls through to the idempotent send).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ed promotion

Round-3 review findings on #516:

- The gate-refused (record-only) path's already-credited guard read only the
  ERPNext Completed row and the intent outcome. Legacy bare-username payments
  have no intent, so a credited-but-not-promoted row re-delivered while ops had
  auto-credit toggled off was stamped `auto-credit-disabled` and paged ops to
  hand-credit money already in the wallet. It now also reads the durable
  credited marker and acks `already_processed`.
- The "credit succeeded but ERPNext promotion failed" alert is now critical.
  The sweep's critical for this state only fires while retry + auto-credit are
  enabled and the row is inside lookbackDays; same dedup key so they fold into
  one incident.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…-out stranded rows; decouple marker TTL from lookback

Review fixes for the stranded-credit sweep:

- treasury-loop: hand the sweep balance - criticalFloorUsd instead of the
  full balance. The floor is a reserve for live card traffic, not a one-shot
  gate; two stranded $280 rows could otherwise drain a $600 float to $40 and
  reproduce the insufficient-treasury-float incident this PR exists to
  prevent. Uncovered alerts now read "available above reserve".
- credited-marker: TTL is max(30 days, lookback + 1) instead of lookback + 1
  read from live config. Raising fygaro.retry.lookbackDays after an incident
  must never expose rows whose markers have already expired to a re-pay.
  Schema comment on lookbackDays states the constraint.
- ErpNext/sweep: candidates are windowed on last_seen_at, which a failed or
  skipped retry never touches, so a row that fails every tick silently left
  the list on day lookback+1. Added countAgedOutUncreditedFygaroTopups (same
  filters, inverted cutoff, shared query helper) and the sweep pages critical
  once per dedup window with the count and oldest request_id, including on
  sweeps with zero in-window candidates.

Tests: treasury-loop reserve hand-off ($600/$500 -> $100), sweep reserve
coverage ($280 row uncovered, not credited), new credited-marker.spec, aged-out
paging (zero candidates, after a sweep, stopped-on-float, count error,
kill-switch), ErpNext aged-out reader (filters, exclusions, oldest, errors),
writer wrapper.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…g payloads for the aged-out count

- The aged-out stranded-topup page told ops to "credit by hand" without
  consulting the Redis credited marker. A row the sweep paid on day 1 whose
  ERPNext promotion failed ages out on day 8 and would be named as the
  oldest row to credit — a second payment. countAgedOutUncreditedFygaroTopups
  now returns the aged-out request_ids (oldest first, capped at 50) and the
  sweep reads the marker for each, splitting the page into "uncredited:
  credit by hand", "already credited but not Completed: promote by hand, do
  NOT re-credit" and "unverified: check before crediting" (marker unreadable
  or beyond the id cap), with the ids in context.
- The `<` (aged-out) path of queryUncreditedFygaroTopups requested the same
  field list as the candidate path, so every refused and email-attributed
  row ever written (all permanently Fiat Received) had its raw_payload_json
  fetched every 15 minutes and discarded JS-side. The count path now asks
  only for request_id, account_id, failure_reason, source_systems_seen,
  last_seen_at.

Tests: aged-out row with credited marker pages promote-not-credit; all-credited
page contains no "credit by hand"; unreadable marker and beyond-cap rows are
unverified; ErpNext count path asserts the trimmed field list and the id cap.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…"credit by hand"

Review round 3 (second pass) on #516. The aged-out set has no lower bound on
last_seen_at while the credited marker lapses after MIN_CREDITED_MARKER_TTL_DAYS.
A row credited on day 1 whose promotion failed and that nobody promoted would be
paged "promote by hand" for 30 days and then flip to "credit by hand" on day 31
when the marker expired — the double-pay trap re-opened past the TTL.

countAgedOutUncreditedFygaroTopups now also returns the capped rows with their
last_seen_at; classifyAgedOut puts any row older than the marker horizon (or with
an unparsable timestamp) straight into `unverified` without reading Redis. Rows
inside the TTL, and id-only readers with no timestamp, keep the marker as
authoritative.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@islandbitcoin
islandbitcoin merged commit dcea51e into main Oct 4, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants