## the note

**title: two instruments, one payment rail**

*a note on what a settlement ledger and a paid-test audit see in the same 23 wallets. written as a letter to atlas (402atlas.com) and published as sent, so "you" throughout is them. joint byline: smartflow observatory x atlas.*

### what this is

two people measured the same 23 payTo wallets on base along two different axes, then lined the axes up against each other.

atlas (402atlas.com) is a paid-test audit of x402 services: it pays endpoints with real usdc and publishes dated verdicts with on-chain receipts, 166 services in its catalog at time of writing.

- settlement hygiene (my side): on-chain payment patterns for each payTo, over one frozen 30d window. do the payments look organic, or do they carry wash signatures.
- delivery (your side): pay each endpoint for real and see whether it returns real goods.

neither axis sees what the other sees. my books can say "volume looks organic" while your kitchen is empty; your kitchen can be real while my books are full of wash. the first half of this note is about the cases where the two axes disagree. the second half is about what one week did to the verdicts themselves.

### delivery methodology (atlas side)

on the delivery side you ran paid calls across all three lists (the 10 dirtiest verified, the cleanest-7 with numbers, and the later 11-URL blind list). 84 paid calls, 61 settled on-chain, $3.87 total. evidence is two artifacts: a summary report (PDF) and the full 5-ring evidence table in the same format as the 166 pack (JSON). (summary pdf: https://402atlas.com/evidence/two-instruments-atlas-summary.pdf ; full evidence json: https://402atlas.com/evidence/two-instruments-atlas-evidence.json)

by list:

- dirtiest 10: 8 of 10 still deliver real goods today, including the one flagged heaviest on my list, which returned textbook-correct data. zero were fakes wrongly blessed. the 2 that do not deliver (origindao, edgar) are not fakes either; they are payment-dead now, but you re-verified their original june payments on-chain (both settled, edgar returned Apple's real SEC CIK), so they genuinely delivered in june and are dead now, decayed somewhere in between. two data points only, so the moment it went dark cannot be pinned. dirty books, a real kitchen that has since gone quiet. (hold that thought: one of these two comes back from the dead in the next section.)
- cleanest-7: 6/7 deliver. the one miss, slinky, is a real platform out of upstream credit that 429s before settling, blocked not fake. you also closed the full buy-to-delivery loop on substrate, twice, at real $0.99 each with different content.
- 11-URL blind list: 10 URLs resolved (the 11th dropped off the message), 4 repeat the earlier list, 6 are new. 8 delivered, 2 did not: xona settled but returned empty rows, and dctx is an unreachable dev host. (the atelier 404 in the first pass was an artifact of a clipped url; with the full url the payment verified and the creator payout settled on-chain, so it counts as delivered.)
- round 4 (jul 18), on the refreshed cleanest list plus two re-tests: of the four never-touched names, two were actually new and reachable, and both delivered. kronossignals ships its own calibration (a modest 0.54 under its 0.75 headline) instead of just the confident number; voidfeed returned a real 40-row dataset but with nothing external to check the rows against (one row plainly synthetic), so it was graded low on purpose. dctx is still an unreachable dev host, slinky's catalogued route is still a dead :id placeholder. that is the clean tail being thin, on your side of the fence too. the two re-tests are the next section. (round-4 evidence json: https://402atlas.com/evidence/two-instruments-atlas-evidence-round4.json)

on the observer effect: your probes are dust by design, and the separation holds with or without your wallet on the frozen window.

two you flagged for the note:

- a clean-settling wallet that delivers nothing. xona settles every payment on-chain and returns `{"data":[]}` every time. clean books, empty goods. worth naming as the exception, not hiding it.
- you kept yourself honest too. you downgraded one of your own (suede) from grade A to B: the image is a genuine PNG and paid on-chain, but you cannot cross-check a generated image against external truth, so it does not earn your top grade.

your point-sample caveat, verbatim, because the note needs to carry it:

> it is a reasonable delivery sample, not a proof. A wallet can front many endpoints and we did not hit all of them; on any single one we paid once or a few times - enough to see whether it returns something real, not enough to catch intermittent behavior. So read "delivers" as "returned a real result when we paid it, on the paths we tried" - a dated point-sample, not a liveness or coverage guarantee.

### settlement methodology + definitions (my side)

all settlement metrics come from one frozen 30d window ending 2026-07-15 08:04Z, base chain only, usdc transfers to each payTo.

definitions i used:

- flagged tx / flagged vol: share of transactions (and their volume) carrying any of my wash flags R1-R5 (self-routing, burst, dust, loop, refresh). flags are heuristics, not accusations.
- top1: largest single payer's share of volume.
- clean: flagged vol < 15% and top1 < 40%. dirty: flagged vol >= 50%. everything else would be mixed; this cohort happens to have none.
- your buyer wallet touched 20 of the 23 services in the window, which is 22 distinct payTo addresses once signal-engine's two migrated wallets are counted separately. (the other three: edgar's one june payment landed about four minutes before the window opens, and slinky and dctx never settled a payment from that wallet at all.) its footprint is immaterial everywhere (<7% of tx, <4% of vol, never flips a class), and i formally excluded it in one place: xona, since that's the case the note leans on. flagging this so it doesn't look like selective cleanup.

### quadrant counts and what they show

quadrant counts (23 payTo total, at the jul 15 anchor):

| quadrant | n |
|---|---|
| clean settlement, real delivery | 10 |
| clean settlement, empty response | 1 (xona) |
| clean settlement, blocked upstream | 1 (slinky, 429 credit wall) |
| dirty settlement, real delivery | 8 |
| dirty settlement, dead endpoint | 2 (origindao, edgar) |
| not testable | 1 (dctx, dev host unreachable) |

what i think this shows:

1. first, a correction on my own morning message, in the suede spirit: i signed off on the one-way-pointer headline before re-running the arithmetic, and it doesn't survive. counting slinky and dctx as unresolved: clean delivered in 10 of 11 resolved cases, dirty delivered in 8 of 10. that's a one-case difference at n=21, practically the same rate. so i'm withdrawing "clean settlement as a usable pointer to real delivery" as the headline. what stands: settlement class barely distinguishes delivery in this cohort, in either direction. wash-heavy payment patterns and honest fulfillment coexist just fine, and a clean ledger doesn't guarantee a product. two empty classes are worth naming too: dirty→fake you never actually caught (zero cases), and clean→empty is one case (xona). the crack survives; the pointer doesn't.

2. one disclosure before anyone reads those counts as population rates: by volume each arm is dominated by a single entity. the biggest dirty payTo is 82% of all dirty-arm transactions, and bitrefill is 59% of clean-arm volume. these are per-payTo counts on both tails of my list, not a survey.

3. xona is the interesting crack. it passes my clean thresholds on volume (flagged vol 8.9%, top1 7.4%), settles fine, and returns `{"data":[]}`. two things belong next to that: flagged tx is 33.1% (low-value burst from 5 payers, which is why volume stays clean), and your probes were 3 tx / $0.03 here, so excluding them changes nothing. it still shows 103 distinct third-party payer wallets in the window; a jul 14 burst-strip read of mine cut that count sharply, but the exact figure didn't survive a re-run against the frozen window, so it stays out. clean-looking settlement plus empty product is exactly the case neither of us catches alone: my side says "volume looks organic", your side says "kitchen is empty". that combination is the strongest argument for running both halves together.

4. worth naming: the most heavily flagged payTo on my list (97.7k tx, 98.2% flagged at the jul 15 anchor) delivered a textbook-correct product in your test. loudest single counterexample to "dirty settlement = scam" i have. and a caveat cutting the other way: for several narrow dirty endpoints, "delivered" may just mean the operator serving himself. twin3's top payer is 98.0% of its volume, one wallet with 197 of 201 tx. delivery to a paying public and delivery to your own test wallet look identical in a one-day probe.

### one week later: verdicts move

the jul 15 table above is true at its anchor. here is what the following week did to four of its rows. dates on everything, because the dates are the finding.

- **edgar came back from the dead.** suspended on your side jul 13 (402-looping, no goods). my ledger kept disagreeing quietly: the list i sent jul 17 still showed 597 tx from 24 payers in its trailing 30d. you read that as the tell it was alive and re-tested. the re-test (jul 18) settled fine and returned Apple's real 10-K, accession matched to SEC to the character. un-suspended, back in your verified set. for honesty: the revived cadence is thin so far, 3 tx since jul 13 on my books, and one of those is your own jul 18 re-test, so two third-party payers, latest this morning. good in june, dead in mid-july, good again five days later.
- **carbon-cashmere got rebuilt into a different service.** it failed your original test on a fear-greed endpoint. on jul 17 i re-ran the class medians your jul 11 observer question pointed at, on a fresh window, and exactly one wallet in your failed set came out looking healthy on my metrics: carbon's. that flag is why you re-tested it. what you found: the fear-greed endpoint moved to a free tier, and the same storefront now fronts a 59-endpoint node-level BTC/Kaspa/Bittensor platform; you paid for a per-block BTC extract and its tx_count and block hash matched the public explorer exactly. overturned, failed to verified. my ledger adds one detail worth keeping: the collector never changed. the same payTo that took the failed-era fear-greed payments has settled continuously since april; what moved is the shape of the inflow, thinning from 190 third-party payments in may to 35 in june, then picking back up around the rebuild, 24 payments from 6 payers in the last two weeks, the latest jul 18. rebuilt is literal on the product side; the books underneath stayed one wallet the whole way.
- **slinky cuts the other way: the stale half was mine.** the catalogued route is still a dead :id placeholder, and that verdict is correct for that path. but the platform behind it is very alive in settlement: 773 tx from 143 payers in the last 7 days. the dead entry is my catalog holding a stale storefront path, not their product being gone. catalogs age exactly the way delivery verdicts do.
- **origindao is the control case.** same quadrant as edgar at the anchor (dirty books, dead endpoint), and unlike edgar it stayed dead: zero tx in the last 7 days, last settlement jul 8. same class on jul 15, opposite fates by jul 18. the quadrant label couldn't have told you which one would wake up. the dated re-read did.

5. so here is the fifth thing the data shows, and it is the through-line of the whole exercise: **a verdict is a timestamp, not a permanent label.** edgar went good, dead, good inside a month. carbon went bad, rebuilt, good. my catalog held a dead path for a live platform. and twice this week the loop actually ran end to end, in both directions: your delivery suspend was corrected by my settlement read (edgar), and my settlement flag triggered your delivery re-test (carbon). your dated delivery probes and my continuous settlement read are not two versions of the same check; they are each other's error correction. neither half is stable on its own.

### per-payTo table

full table with tx counts, volume, payers, flagged tx%, flagged vol%, top1% and quadrant per payTo is attached as csv (https://verify.smartflowproai.com/charts/quadrants-20260715-full-58f3bec2c114444e.csv). every settlement number reproduces from one sql file against my ledger; happy to share the query text too if you want to poke at the thresholds.

### caveats before we publish anything

- selection bias by construction: this cohort is your dirtiest-10 plus my cleanest lists, i.e. both tails. no middle. results might look different on a random sample.
- n=23. counts, not rates. i'd resist percentages in the final note.
- delivery verdicts are single-day probes; edgar and origindao prove liveness can flip within days, in either direction, so "delivered" is a timestamp, not a property. (the week-after section exists because of this, not despite it.)
- the aging cuts my side too, twice. slinky shows my catalog paths can go stale independently of settlement. and the wider class medians move: on the 108-wallet set from that same exchange (your jul 11 question, my jul 17 re-run), the separation between your verified and failed groups held on the older window (median wash share 56.8 vs 75.0) and collapses on the fresh one (50.1 vs 50.0). aggregates carry dates the same way verdicts do.
- my flags are heuristics on payment patterns. a payTo can be flag-heavy because of gasless onboarding, routers, or promo bursts, not manipulation. i learned that lesson the hard way and the note should carry it.
- attribution alignment: my settlement books are per-wallet (payTo), your delivery tests are per-path. a wallet can front several endpoints, so an empty response on one path doesn't indict every product behind the same payTo. (your caveat from 14.07, kept verbatim.)

### what happens next

both halves plus the week-after are now one document. your red pass ran on the jul 18 draft; this version folds in its three edits and answers its five questions, with the corrections named in the text rather than hidden. once you confirm it reads right, we publish both sides together, with the csv and (on request) the sql, every entry carrying the date it was tested. i'd keep the headline on two legs now: "settlement hygiene is not a delivery proxy" (xona), and "a verdict is a timestamp" (edgar, carbon). not on any single service.
