A generative AI does not return a stored result: it writes the answer each time, with some randomness, with that day’s index and, on some platforms, with your location or session. So the same question can produce different lists on Monday and Tuesday.
That does not invalidate measurement; it makes repetition compulsory. Each combination of question and platform is labelled across rounds: present in all, absent in all, variable or not assessable. A stable overall rate can hide the fact that almost every row changed.