How long do you trust an A/B test result before you re-check it?

Something I keep running into and I still don’t have a rule for it.

You test a setup, one variant wins, you ship it and move on. Standard. But that test ran on last year’s traffic mix, in a different season, before AI assistants were sending anybody to storefronts at all. I have no idea whether the winner is still the winner. Usually I only find out because some number drifted and I went looking for why.

Where I’m coming from, we build onsite popup and signup widgets, so I get to look at a lot of stores’ capture setups. Almost none of them have a date attached to why the current version is the current version. Somebody tested it at some point and it stuck.

Three things that feel like they should force a re-check.

Season. A test that won in July was measured on July intent. BFCM traffic behaves nothing like summer browsing, and nobody re-runs the July test in November.

Traffic mix. If paid went from a fifth of your sessions to half of them, the audience that variant won on isn’t the audience you have now. Same story when a channel dries up.

AI referrals. This is the one I’m least confident about. The share of sessions coming in from LLM assistants keeps climbing, and those visitors show up pre-researched with one specific product in mind. A popup tuned to interrupt somebody who’s browsing is now being shown to somebody who came to buy one thing. It seems like that should change which variant wins, but I can’t prove it yet.

Rollouts being in early access makes this more pressing rather than less. A lot more stores are about to be testing natively, which means a lot more results piling up, and nobody goes back to check the old ones.

So, the actual question. Does anybody here have a written rule for re-testing? Something like re-check every winner after two quarters, or re-check when any channel shifts more than some percentage of sessions. And if you have one, what fires it, a calendar reminder or a metric threshold?

If you’ve got one that works I’ll take it.

My rule is 90 days for high-impact tests, 180 days for everything else. I also re-test early if any of these happen:

  • Traffic mix shifts by 20% or more for a major channel.
  • Mobile vs desktop mix changes by 15% or more.
  • A major season starts, especially BFCM, holiday gifting, or a store-specific peak.
  • The winning metric drops outside its normal 4-week range for 2 straight weeks.

I keep a simple log with test dates, traffic split, season, winner, and lift. Every month I sort by oldest decision date and pick one stale winner to challenge.

For AI referrals, I would not make a separate variant until that segment can produce a properly powered test. Today, I’d at least exclude obvious high-intent product landing sessions from browse-focused popups, then compare capture and purchase rates by source.

Good thresholds. I’d expire decisions by mechanism, not by calendar alone. A copy-clarity win may survive longer than an urgency popup or an acquisition-specific offer.

Store the test’s traffic, device, season and offer context, then monitor the original primary metric plus its guardrail. Retest when the exposure context changes enough that the original causal story may no longer hold. A winner is evidence for a context, not a permanent property of the page.