Polluted research is poisoning AI tools

AI chatbots have a well-known problem with misinformation, often providing incorrect information, fabricating citations, and struggling with math. A common assumption is that this issue will resolve itself as AI models grow and their infrastructure improves. However, as AI becomes increasingly embedded in Canadian society, a growing pipeline of politically manipulated research is influencing the outputs of AI tools.
The Erosion of Scientific Publishing
The scientific publishing ecosystem has been deteriorating for decades, with the proliferation of predatory and pay-to-play journals, as well as weak or ideologically motivated studies being presented as legitimate research. Vass Bednar, managing director of the Canadian SHIELD Institute, notes that the assumption that published research is peer-reviewed is no longer valid, and that journals are becoming increasingly “slop-ified.”
Large language models are not designed to determine truth, but rather to detect patterns in vast amounts of text, including opinion articles, social media comments, and government reports. Thor Tronrud, senior machine learning scientist at StarFish Medical, explains that during pre-training, every piece of text is treated equally, with fiction being equivalent to fact and highly peer-reviewed research.
Read Also: Canadian Hospitals Debate Security vs. Design to Curb Violence
The Vulnerability of AI Models
Companies attempt to filter out low-quality training data using automated heuristics, but these methods are not foolproof. Tronrud notes that the quality of the training data is only as good as the signals used to assess it, and that the need for vast amounts of text can lead to a tradeoff in quality. This tradeoff presents a new national vulnerability for Canada, as politically aligned actors in the United States are constructing a parallel ecosystem of academic journals and research bodies that can be mistaken for credible sources by AI models.
Government-affiliated or institutionally branded content receives strong weighting during the training process, which can be problematic as science institutions fall to political capture. Dr. Peter Hotez, Dean of the National School of Tropical Medicine at Baylor College of Medicine, warns that U.S. government officials are assembling “a whole alternative universe of pseudoscience” complete with journals and institutional trappings.
The Consequences of Polluted Inputs
Once ingested, pseudoscientific material can become indistinguishable from legitimate research within AI models, leading to the presentation of politicized claims as evidence. Retractions of flawed studies do not “unbake” a model’s internal weights, allowing junk science to be summarized and effectively laundered by the model. This can have serious consequences, as Bednar notes that research constructed to promote bogus claims can become repeated and cited elsewhere, making it harder to fight.
