What it measures
Two independent signals over the complete endpoints catalog: a payout destination declared under more than one host, and the same path, response-shape hash, and payment-required amount appearing under more than one host.
How it is produced
The canonical read-only script is version two of the duplicate-host measurement. It normalizes URL netloc values to lowercase, aggregates both hostnames and registrable domains, excludes unresolved URL templates, parses declared payout destinations from stored response payloads, and separately builds content-and-price signatures. It reports parsing failures and truncation rather than treating missing payload data as evidence of absence.
Canonical source
duplicate-hosts-measure-v2.py
Ad hoc queries and the deprecated version-one script are not cited sources for this measure.
Limits and assumptions
- A shared payout destination is not the same as a shared operator. One platform can serve many independent operators, so Signal A is an upper bound on the phenomenon, not a count of false services.
- Signal B also captures legitimate multi-domain deployments and white-label services.
- The denominator is the full catalog and is not directly comparable with a paid-tested list. The populations must not be presented as if they were the same.
- Payout-destination coverage must be reported against parseable payload rows with a separate truncation count, never as coverage of the entire catalog. Signal A cannot see routes without a declared destination or rows truncated by the crawler.
- The database changes over time, so every citation requires the run date.
- Registrable-domain aggregation depends on the public-suffix data used by the script, including its handling of private hosting suffixes.
How to refute this
- Show that the largest Signal A cluster is one legitimate platform, making the aggregate a platform artifact rather than duplicate services. Always report the cluster-size distribution and name the concentration when publication is allowed.
- Run Signal A with the largest cluster excluded before publishing a result.
- Show that Signal B measures shared framework templates rather than duplicates. Review the top collisions manually; exclude a common SDK signature or lower the claim if it joins unrelated operators.
- Show that Signal B drifts by more than 10% between canonical runs one day apart. Before publication, use two runs at least 24 hours apart and cite both or the window median.
Last validated
2026-08-30
Run timestamp 13:58:15Z. This measure is citable only with its run time, because the catalog changes between runs.