A brand new preprint examined whether or not tweaking only one a part of a supply modifications AI search citations, with all the pieces else stored the identical. Within the uncooked numbers, the highest consequence received cited about twice as usually because the fifth. When researchers Sriram Selvam and Anneswa Ghosh reversed the order of matched sources, the influence was a lot smaller and measured zero in a follow-up check.
Posted to arXiv on September 14, the paper isn’t peer-reviewed. The research covers one GPT-5.4 search agent that makes use of Exa as its search supplier, that includes offline replayed conversations and no dwell webpage edits.
How The Take a look at Labored
The researchers prompted the GPT-5.4 agent to reply 130 widespread questions by having it carry out unbiased net searches. They recorded each message and search consequence from the 129 questions it addressed. From these transcripts, they selected pairs of pages that appeared in the identical search outcomes and had been each screened as supporting the identical truth. This screening aimed to seek out conditions the place both web page might be pretty cited. When true matches had been recognized, any credit score variations had been because of how the mannequin apportioned recognition between the 2 sources, each confirming the identical truth.
That left 113 pairs. A later blinded human verify confirmed 103 of them as real matches. The researchers replayed every saved dialog 4 methods, putting one web page above or beneath the opposite and exhibiting its textual content both as plain paragraphs or rewritten with headings and lists or a desk. Solely the ultimate reply was generated once more.
Each variations of the textual content had been generated by AI rewrites of the unique web page. Grok 4.3 created almost all of them, with GPT-5.4 used as a fallback for one pair, and a separate Grok evaluation checked that the information matched. The wording varies between the 2 variations, so the authors notice that the check compares two rewrites however doesn’t particularly isolate formatting variations.
Uncooked Place Hole Was Bigger Than Swap Results
Within the preliminary place of a search name, pages had been cited 85.1% of the time in saved transcripts, in comparison with 42.8% for pages within the fifth place. This creates a distinction of 42.3 share factors.
Right here, ‘place’ merely means the order of the 5 Exa outcomes returned in a single search, not the place a web page ranks on Google or its place on the dwell net.
The research highlights that search suppliers often put extra related pages on the high, so the uncooked distinction displays each the place and the standard of the pages. When the researchers moved the identical web page greater inside its pair, the prospect it was cited in any respect went up by 7.9 share factors. Nevertheless, this discovering wasn’t thought-about statistically important after accounting for a number of exams.
One other testing set with 56 pairs, the place solely the order was switched, confirmed an estimate of 0.0 factors, with a 95% confidence interval from -5.4 to +5.4.
The research explains that the uncooked hole and the swap outcomes measure totally different elements. General, it means that place did affect citations in some instances, however averages from uncooked place knowledge aren’t dependable.
Structured Rewrites Bought Extra Credit score, Not Clearer Entry
Pages that had been rewritten with headings and lists acquired a mean of 0.50 extra quotation markers per reply in comparison with the identical pages written as plain paragraphs, with a 95% confidence interval from 0.20 to 0.84. The solutions within the check had been closely cited, with a median of 29 markers throughout six paperwork.
The full variety of citations per reply didn’t rise, and the quantity on the opposite web page barely modified. The authors see this as credit score being targeted extra on the rewritten web page.
The principle check the researchers performed, which they deliberate earlier than beginning the experiment, was to see if the web page received cited in any respect. They discovered that utilizing structured textual content elevated that chance by 4.5 share factors, with a 95% interval from -1.4 to +10.4. The paper factors out that this consequence isn’t conclusive and mentions that the research may reliably detect solely results of about 8.5 factors or extra.
A extra strict comparability, the place each phrase stayed the identical however the structure was adjusted to at least one sentence per record row, boosted quotation charges throughout all 113 pairs. After they repeated the check with a subset, the impact reversed.
Within the dialogue part of the paper, the authors shared these insights:
“That is an attribution-sensitivity warning, not an optimization tactic.”
Reruns Modified Quotation Outcomes
The researchers examined 120 responses once more utilizing the identical inputs, and located that the choice to quote or not for the goal web page modified in 15% of those instances, roughly one in seven.
The typical depend impact remained constant throughout these reruns. They estimate that about 45% of the variation in a single run’s impact is because of mannequin randomness.
The authors advocate rerunning quotation exams a number of occasions and sharing how constant the outcomes are throughout these runs.
Moreover, SparkToro reported in January that ChatGPT and Google’s AI Overviews every produced the identical model record lower than 1% of the time when given the identical immediate repeatedly.
Why This Issues
The uncooked place hole on this check was a lot bigger than the typical impact noticed when researchers swapped supply order. An Ahrefs report from Could confirmed pages cited by AI had been about 3 times extra more likely to embrace JSON-LD schema, however including schema didn’t clearly enhance citations.
This raises questions on whether or not a correlation in a vendor report or your monitoring was ever examined by altering the variable, and a single reply is a weak foundation for labeling a quotation as gained or misplaced.
The research can’t verify if reformatting a dwell web page boosts citations, since rewrites solely utilized to textual content already retrieved, excluding crawling, retrieval, and rating processes.
Wanting Forward
The researchers re-ran the saved searches on Grok 4.3, discovering that the structured rewrites leaned the identical manner. Nevertheless, lower than half of Grok’s first replies adopted the proper quotation format.
The authors advocate extra analysis to check every state of affairs a couple of occasions, discover totally different search suppliers and fashions, and take note of each how usually citations happen and if a web page is cited in any respect.
Featured Picture: Accogliente Design/Shutterstock
