A legally clean alternative to platform review data: every metric here comes from pages destinations publish about themselves. 43 sites resolved to 33 domains, measured live and reconstructed year‑by‑year from web archives, 2022‑2026. Every coded variable has been re‑measured by three independent LLM coders; one headline finding did not survive that check.
Borobudur/Prambanan, Besakih, Lake Toba, Komodo and Raja Ampat had no official site in 2022. All five are corroborated by domain registration dates, so this is not an archive-coverage artefact. Montenegro started saturated at 94% and added none. Coded from captures and RDAP, not from the failed instrument — unaffected by the correction above.
Fisher exact · p = 0.018 · survives FDR · standsWithdrawn under the pre-registered rule (κ < 0.60 → the finding does not stand, whatever its p-value), then re-measured across all 142 archived domain-years by a coder three independent models agree with. The direction survives — Indonesia is higher in every one of the five years — but the difference is not significant at n = 14 vs 16, and neither country shows any adoption trend.
Fisher exact · p = 0.204 · withdrawn, then not confirmedMarginal, but Montenegro is ahead in every single year of the panel (56–69% vs 22–38%). The gap is stable, not a one-year artefact. Reported as a directional tendency, never as a finding.
Fisher exact · p = 0.063 · not significantOne finding stands. The two-directional story this dashboard told earlier — Indonesia ahead on transacting, Montenegro ahead on being found — rested half on a measure that turned out not to measure what it claimed. What survives is narrower and firmer: Indonesia is building presence where Montenegro was already saturated, and the discoverability gap runs the other way but does not clear significance. A single-year cross-section of the same 43 sites found no significant difference on any technical feature.
The instrument failure is itself a result. It is arguably the most portable thing here: keyword-and-platform detection is a common instrument in web-audit studies of tourism websites, it is almost never accompanied by a reliability figure, and on this multilingual corpus it performs at chance.
Every variable re-coded blind by four coders — the original instrument plus three LLMs from three providers, at different scales and on different infrastructure. Cohen's κ, 33 domains, all six pairs.
| Coder | Model | Where it ran |
|---|
Model identity is taken from the gateway's configuration, not from the models' own statements — and that is not pedantry here, because both hosted coders misidentified themselves when asked. One answered “I'm Claude, made by Anthropic”; the other “Built by OpenAI, GPT-4”. Since coder independence is the entire basis of this check, a study that establishes it by asking the coders who they are has established nothing.
Read the V1 block as a diagnosis, not a disappointment. If the construct “can a visitor transact online” were genuinely ambiguous, the three LLM coders would disagree with each other too. They do not — they reach κ = 0.897 to 1.000, and a 27B open-weights model running on a laptop reproduces a frontier model's judgement exactly. Every one of them agrees with the regex instrument at chance. The construct is highly codeable; the instrument was not a valid measure of it.
Where the instrument went wrong — three error classes, all visible in the disagreement set. Missed transactional language: “Elektronske ulaznice” (Montenegrin e-tickets) on nparkovi.me, “Beli TLPJL” on Raja Ampat's site, the eRinjani booking system. Keyword without a route: a “Reserve” menu item, the word “reservation” inside a policy, a “checkout” string in a script — six domains scored 1 on that basis and 0 by all three LLM coders. Absence read as evidence: where a page delivered no text, the instrument recorded “no booking”; it had no “cannot determine” category at all, and 28 of 142 panel cells belong in one.
Identical underlying pages. The substantive conclusion depends on who coded them — which is the point, and the reason the reliability check came before the statistics.
| Instrument | Indonesia | Montenegro | Fisher p |
|---|
Live cross-section, 33 domains, 4 September 2026. The three valid coders find a larger gap than the broken instrument did, in the same direction, at p ≈ 0.06 on a single year. So after the withdrawal the hypothesis was unmeasured, not refuted — which is what made re-running the full panel worth doing.
| Instrument | Indonesia | Montenegro | Fisher p |
|---|
Panel, 142 archived domain-years, all re-coded by Qwen3.8-27B locally; 142/142 parsed, zero failures. The regex instrument roughly doubled both rates and inflated the gap. Note that the direction of the correction reverses between the two datasets — larger gap on the live pages, smaller across the panel. An invalid instrument does not err consistently, which is precisely why its output cannot be trusted even when it happens to point the expected way.
Neither country is adopting. Indonesia sits higher in every year under the valid instrument, but online booking capability on official destination sites is not spreading over 2022–2026 in either system. That flatness runs against the assumption that destination e-commerce steadily diffuses, and is worth reporting in its own right.
What may not be claimed. Three models agreeing establishes consistency, not accuracy — they may share a bias. That the LLM coding is correct and the regex wrong needs a human-coded subsample, which has not been run; adjudication_worksheet.json is prepared for it (~26 domains). What holds without any ground truth is the pair of measured facts above: the instruments disagree at chance, and the conclusion flips depending on which one is used.
Three pre-specified hypotheses, corrected as a family. Everything else run in this project is exploratory and carries no claim.
| Hypothesis | Indonesia | Montenegro | Raw p | Bonferroni | BH (FDR) | Status |
|---|---|---|---|---|---|---|
| H1 · new web presence | 5/16 | 0/17 | 0.018 | fails | survives | stands fragile · 1 change → 0.085 |
| H2 · booking / e-ticket | 5/14 | 2/16 | 0.204 | n/a | n/a | withdrawn, then not confirmed |
| H3 · schema.org | 5/14 | 12/16 | 0.063 | fails | fails | directional tendency |
One of three carries a claim. H2's row shows the re-measured figures, not the withdrawn ones; no multiple-comparison correction applies to it because it no longer enters the family as a positive result. H1 survives Benjamini–Hochberg, the pre-registered rule, and fails Bonferroni — which is stated here rather than omitted.
Reliability precedes significance. The pre-registered rule, written before any coder ran: if κ for the booking variable falls below 0.60, the finding does not stand whatever its p-value. It came in at 0.043. That rule is why a p = 0.026 result was withdrawn rather than defended, and it is the part of this project most worth copying.
H1 remains fragile. A single reclassification — one presence date read differently — pushes it to p = 0.085. Expanding the frame toward ~30 domains per country is the remedy, and on that sample H1 and H3 are genuinely pre-specified by 04_analysis_prespecification.md.
The archive-coverage confound was tested, not assumed. Montenegro is archived more densely than Indonesia (median 46 vs 32 snapshots), so an absent capture could have meant an unarchived site rather than no site. Domain registration dates settle it from outside the archive: all four Indonesian domains the archive dates to 2023–24 register in the same window, and Raja Ampat’s kkprajaampat.id — live, serving HTTP 200, registered 2026-05-16, not yet crawled — is a fifth case the archive alone would have missed.
Share of each country's domains with any archived capture that year.
Montenegro's single decline is kotor.travel — the only domain in either country to lose presence. It is still live; it migrated to a client-rendered single-page app and became invisible to archival crawlers. Destination digital heritage is being lost silently.
Percentage of that year's valid captures carrying each feature. Denominators shift as sites appear and as captures fail — the n is labelled on hover. Booking is not shown here: it appears in its own section above, under both instruments, because the regex series is not a valid measurement.
HTTPS is shown for completeness but Montenegro's apparent decline is an artefact: the measure is the URL scheme the crawler happened to archive, and a site reachable over both schemes can be archived under either. Do not report it as a finding.
Not every archived capture is a usable measurement. Screening happens before analysis, never after.
Six captures returned exactly 73 characters under the title “One moment, please…”. Without this screen they would have entered the content series as near-zero, and a site that started blocking crawlers would read as a site that lost its content. 18 of 142 captures are interstitials or empty shells, and the share is rising in both countries.
The live audit, 4 September 2026. All 43 sites in the sampling frame, including those with no official website at all. The Booking column is the original regex instrument's output, kept visible so its errors can be inspected — it is not the measurement any finding rests on.
| ID | Site | Position | Publisher | No-JS delivery | Schema | Booking (regex — invalid) | Analytics | Langs |
|---|