An honest, bounded event-contrast of daily footfall at a single anonymized airport-transit duty-free outlet (Site B) around a major international sporting event. The headline finding: no statistically significant full-store footfall lift when the pre-baseline is properly widened (DOW-adjusted −2,170/day, p=0.458) — the window contrast is decided by the pre-baseline choice, not the event itself. A permutation placebo null with 1,486 random boundary placements and a pre-baseline-robustness battery (three specifications) confirm the result is specification-dependent. At the within-window level, a mid-window transitory hump with full post-window reversion is visible. Includes full confound-and-limitations disclosure. All figures and numbers are reproducible from frozen data. Aggregated, anonymized, relative volumes.
Download the full paper (PDF) →A fail-closed validation-and-guard scheme for demand-forecasting pipelines whose per-entity signals are often data-scarce. Couples a data-reality sufficiency gate (non-zero target density), a serve-time degenerate/flat guard (fail-closed on flat signature or smaPE >95, deliberately not >50 so bias-class forecasts keep serving), a promotion/activation guard (broken retrains cannot displace healthy siblings), healthy-cohort regression + delta deploy-blocking gates, and bias-class vs. flat-class separation. Three addenda: a production training-sufficiency gate, a validation replay across an extended horizon, and an empirical coverage-calibration study of the served prediction band. Observed, aggregated, anonymized; all figures injected from output.json.
Download the full paper (PDF) →The first defensible retail-visitation insight from observed, reproducible data: footfall is powerfully concentrated in time. Across 5,697 complete location-days at 186 cohort locations (open-hours capture, local timezone), the single busiest hour carries on average 21.8% of a day's visitors and the top three hours 53.1%; the modal open window is 10:00–18:00 local (median 9h); weekends run ~1.79× weekdays (median 1.56×); and on a size-invariant per-store index Saturday is the peak day (1.50× a store's average) with Tuesday the trough (0.83×). Every figure is an observed entrance-crossing measurement, never a forecast, and is reproducible from the project's analysis script and results file — aggregated, anonymized, with the N=186-cohort scoping and frozen-snapshot caveats stated.
Download the full paper (PDF) →Is footfall really changing, or is it the weather? Using a weather/calendar ARX compensator on 24 well-covered stores (81-day window), compensation removes ~66% of day-to-day footfall variance at the median (100% of stores improve). A pronounced mid-summer raw-series boom-and-bust flattens into a flat compensated trend, meaning much of the apparent late-summer decline is weather/calendar-attributable, not a genuine demand collapse. Calendar seasonality dominates (44% of daily variance) vs weather (10%). All figures observed, aggregated, and anonymized; caveats preserved (small cohort, descriptive, no significance claims).
Download the full paper (PDF) →Over 20% of a retail location's daily footfall lands in the single busiest hour, and the week peaks on a Saturday, not a weekday. Across 5,697 complete store-days at 186 stores, the single busiest hour carries on average 21.8% of that day's visitors (top three hours 53.1%), with a modal peak at 15:00 on weekdays and 14:00 on weekends (local time); weekends average ~1.79× weekdays, and Saturday indexes at 1.50× vs a Tuesday trough of 0.83×. Observed entrance-line data, aggregated and anonymized.
Download the full paper (PDF) →86% of well-covered stores are weekend-driven (n=71), but the surge is format-specific: standard 9-hour stores are 100% weekend-driven at 2.40× weekdays, while long-window stores average 1.28×. The most time-compressed stores are the small ones — they pack ~63% of a day into three hours. Read together, the busy period is a sharp, format-specific Saturday peak. All figures observed, aggregated, and anonymized; caveats preserved (cohort subsets, no significance claims).
Download the full paper (PDF) →A store facing a busy street is not necessarily a busy store. Only ~0.47% of passers-by walk into the one location with both counters; entering traffic is sharply concentrated (peak hour 15:00 local carries 13.2% of a weekday's entrants, 14.2% on weekends), and 87.1% of stores are weekend-driven. A busy walkway is not a busy store. All figures observed, expressed as shares/ratios, aggregated and anonymized with honest n=1 limitation scoped.
Download the full paper (PDF) →A rigorous, reproducible before/after estimator with honest error bars: a weather+calendar counterfactual baseline, a placebo/permutation null (2000 change-point shifts) to establish the noise floor, and a minimum-detectable-effect (MDE) bar. The first demonstration is an honest null — the observed −4.3% swing is below the detectable threshold — with a positive-control check proving the estimator isn't blind. The measurement discipline, not the headline, is the finding. Observed, aggregated, anonymized.
Download the full methodology paper (PDF) →Formalises the three preprocessing gates that govern per-location demand intake — capture-window empirical recovery, entrance-zone aggregation (location_visits is the only lawful total), and genuine-signal scaling (absent zeros are absence, not traffic). The intake-side companion that governs how raw counter data becomes lawful, analyzable demand.
Download the full paper (PDF) →The weekend-peak assumption fails in a transit-adjacent format. Across 27 complete days in August 2026, the highest-demand day of the week is Thursday (≈2,762 location-visits/hr) — 18.1% above Saturday and 16.1% above Sunday, with Monday the trough — and the month's three strongest days are all Thursdays. The rhythm follows scheduled air service, not the retail calendar. Observed, aggregated, anonymized; case-study scope stated.
Download the full paper (PDF) →The six busiest clock hours carry 41.7% of a day's footfall in a large-format transit-adjacent location — and the single busiest hour is 00:00, not a trading hour (≈4,516 location-visits/hr, nearly 1.9× the overall mean). The conventional daytime block (10:00–17:00) accounts for only 29.2% of the day's movement. A transit-adjacent store operates on a clock that looks nothing like a high street. Observed, aggregated, anonymized; case-study scope stated.
Download the full paper (PDF) →A validated catalog of 14 data-integrity checks for retail-visitation pipelines — all-zero/flat feeds, missing-calendar anomalies, gross-over-prediction/flat-serving, DAY-vs-HOUR reconciliation — each grounded in a real platform incident, with an honest fail-closed verdict scheme (fires / resolved-deploy-verified / exposed / pending / not-meaningful; no fabricated coverage).
Download the full paper (PDF) →First causal estimate of payday footfall effects from the explainable-hybrid program: pooled +28.3% [14.5, 43.8] across SE+CH (n=12,214 store-days, 143 stores, store+DOW FE, cluster-robust). Per-market: SE +39.2% [27.8, 51.6]; CH +20.2% [−7.0, 55.3]. Includes pre-registered sensitivity table with honestly disclosed ±1d sign divergence — Phase 2 completed, pre-registered (see the Phase-2 supplement). Aggregated, anonymized, methodology attached.
Download the full paper (PDF) →Resolves the one credibility gap from the Phase-1 analysis. W2.1: holiday-adjacency leave-out resolves the disclosed CH ±1d sign divergence (−33.9% → +15.5%, CI [−6.3, 42.4] incl 0); the divergence was the 2026-08-01 Swiss holiday inside the July payday window, not anchor misspecification. W2.2: holiday days ≈ near-zero closure days (pooled −96.8%), post-holiday null everywhere, NL honestly not-estimable. Headline unchanged: pooled +28.3% [14.5, 43.8] SE+CH. Caveats: the CH reversal rests on a single holiday cluster (2026-08-01) in a 3-cycle panel — cited with the conditional-estimand and fragility caveats, never as a stand-alone positive; post-holiday nulls are "no detectable effect at achievable power," not evidence of absence. Pre-registration trail: SHA-256 2f7081cb (spec served). Aggregated, anonymized, relative effects only.
Download the Phase-2 supplement (PDF) →Frozen pre-registration spec for the explainable-hybrid program's Phase 2, registered before any estimation runs. W2.1 resolves the ±1-day window sign divergence via holiday-adjacency leave-out; W2.2 adds a decomposed holiday/calendar response read as bias/control context. Spec SHA-256: 2f7081cb (file excluding the self-referential SHA line). Aggregated, anonymized, methodology attached.
Frozen pre-registration spec for the zero-shot time-series foundation model (TSFM) benchmark, registered before any Phase-2 estimation runs. Registers the gated served-pair comparison (Chronos-Bolt-small / TimesFM-2.x-small / Moirai-small, CPU-only, zero-config) versus the as-deployed served stack, plus a non-gating deep-subset zone read. Falsifiable decision rule: a MATERIAL ZERO-SHOT WIN needs Δ ≥ 2.0 SMAPE points, a ≥5% relative pooled-typical median win, failure-mode non-regression, and a low-volume-tercile win — otherwise PARITY (served stack favored by parsimony); "served stack wins" is an explicitly valid outcome. Spec SHA-256: 5ac1f9b9 (file excluding the self-referential SHA-256 line). Aggregated, anonymized, methodology attached.
What happens when a pre-registered benchmark rule blocks an attractive accuracy win? We benchmark three 2026-class zero-shot time-series foundation models (Chronos-Bolt-small, TimesFM-2.x-small, Moirai-small) against the probabilistic footfall forecaster our platform deploys in production, on 20,091 same-day band forecasts across 133 benchmark-eligible stores over 3 ISO weeks. Two challengers beat the served stack materially on accuracy (served median SMAPE 44.23 vs 34.11 and 32.36), but the frozen, public decision rule (sha256 5ac1f9b9...) also requires failure-mode non-regression — and both winners fail it. Recorded verdict: "accuracy-only win, disqualified by chronos (failed cond3_failure_nonregression); timesfm (failed cond3_failure_nonregression)". Honest outcome: no TSFM promoted; secondary non-gating deep-subset read finds seasonal-naive more accurate than every zero-shot TSFM. All numbers generated by code from frozen artifacts; no serving decision is made here. Pre-registration spec served publicly.
Download the full paper (PDF) →On a retail estate, a store's footfall does not scale with how many stores its country has. Across an anonymized 201-location, 16-market estate (current 2026-10-01 forecast vintage, week of 1–7 October), one of the largest stores networks — 49 stores, 24% of the footprint — is forecast to carry only ~0.7% of the week's planned trips (~0.03x the per-store average), while a six-store Iberian market is forecast at ~7.7x, contributing ~23%; the single largest network (76 stores, 38% of the footprint) sits almost exactly on the average (~1.03x). The per-store spread between those two markets is ~258x under either benchmark. Caveats are disclosed, not hidden: the figures are a weekly forecast (planned visitor-trips) re-read each forecaster run — not observed counts — single-week point estimates, and the three tail markets are single-venue figures. Estate- and market-aggregated, country-level only; no venue-level data is published.
Download the full paper (PDF) →