How I would know I'm wrong
eBay's randomized shutdown of brand search found conventional attribution implying 4,100% ROI where the experiment measured negative 63%. Any long-window architecture that does not state its falsification criteria up front is selling something.
27 July 2026 · Note
Long-window attribution is where architectures of this class overclaim, and the correction from experimental economics is severe enough that it should change how anyone in this category talks.
Blake, Nosko and Tadelis ran a randomized shutdown of eBay’s brand-keyword search. Ninety-nine and a half percent of paid clicks were substituted by organic traffic. Conventional attribution implied an ROI above 4,100%. The experiment measured negative 63%. The mechanism is not subtle: last-click methods credit the channel with purchases by frequent buyers who were going to transact anyway.
Anyone proposing an intent-continuity architecture is proposing something with a measurement profile at least as bad, because the window is longer and the touchpoints are more numerous.
Endpoints by latency, not by convenience
The honest structure is to sort what you can measure by how long it takes to become measurable, and to refuse to report on the rest.
Days. Cost per zero-party capture. This is the primary endpoint of a first cycle because it is the only one that resolves fast enough to steer anything.
Weeks. Message-to-surface engagement, and token return-visit rate. The second of those is the direct test of the identity claim: if the email link is not re-establishing identity, the architecture’s load-bearing assumption is false.
One full window, minimum. Deposit attribution, and credible only against a holdout rather than against last-click counts.
The three ways this dies
Stating these before the pilot rather than after is the entire discipline.
If cost per zero-party capture is no better than the broad-creative control, specificity does not drive response, and the generation stage is expensive decoration.
If token return rate is indistinguishable from baseline direct traffic, the email link is not functioning as the re-identification mechanism, and the central claim fails.
If validation-gate rejection runs high enough that generated coverage falls materially below the inventory bucket, the asset library’s tagging cannot support the assembly step, and there is no input.
Why publish this
Because the alternative is a category where every proposal works, which is how you can tell none of them are being measured.
Cut from: Carrying creative semantics across the travel consideration window

