Ridley Liu has compiled a sizable list of sources regarding the impacts of generative AI, and written a summary of each source. While this webpage is pretty large, it is by no means the definitive source on the vast reach of Generative AI impacts. This webpage will use the shorthand “AI” to refer generally to large language models and diffusion models used at an industrial scale. While I personally have an issue with calling these programs intelligent (even with the “artificial” designation), this is the term that is most culturally accurate.
The intention of this document is to educate individuals on the common issues people may have with generative/agentic AI.
Ridley Liu
”While my writing here is personal and opinionated, it is crucial to note that the following sources and summaries were written to be objective and present facts about this subject. As of September, 2026, this information can accurately educate a curious reader on the impacts of generative AI.
Of the sources presented, the psychological impacts stood out to me the most. In a world where we all seem to be slipping more and more into echo-chambers, having programs that confidently affirm the sentiments that are written to them will only exacerbate social siloing. It’s a lonely world that’s being built here on the internet. This is why I suggest reaching out to family, friends, and loved ones and fight back against the loneliness that is the social winter we find ourselves in.
A warning: The topics covered on this page include sensitive subjects such as pornography, CSAM, suicide, sexual harassment, and grooming. Proceed at your own discretion.
The Common Crawl is a database that has been around for years that scrapes the entirety of the internet, including paywalled articles, copyrighted material, medical documents, and anything else that could possibly be available on the internet. This affects the outputs of said models in ways that no one (including the creators) can fully predict or control. There is no way to source the amount of data needed to train a model as large as ChatGPT, Claude, Grok, etc. without scraping private material.
Cara is a site created for artists to avoid having their work scraped by AI. When there are spaces explicitly created to escape AI, people take it upon themselves to find ways to scrape those sites. A clear statement expressing a lack of consent does not prevent scraping, and sometimes encourages it.
Amazon was exposed in an investigation for buying tons of rare books (“A rare book is most likely a book that was never printed in large quantities because it is about a niche subject.”) and physically cutting them apart to scan them and feed the content into their models. There have been other many documented cases of AI companies buying books en masse and destroying the physical copies in the process of scanning the content. This is particularly harmful when they are destroying books that are not widely available.
We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear.
Shumailov, I., Shumaylov, Z., Zhao, Y. et al.
”Model collapse is when models start to consume content produced by other models rather than human made content. This causes the following outputs to degrade in quality. “We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear.” (A quote from the study linked below done in 2024, updated in 2025). As more and more content is generated, the harder it is for models to avoid consuming generated content, especially since these models rely on mass web scrapers.
AI models collapse when trained on recursively generated data
Once an AI has been trained, there is no way to retroactively remove any select training data from it. The only way to “remove” that influence completely is by entirely retraining the model without that data. Retraining a model entirely from scratch is a very expensive endeavor and not done frequently. This means copyrighted material, sensitive information, and other AI generated content, cannot be reliably scrubbed from any data set that is being used to train large models.
Does your LLM truly unlearn? An embarrassingly simple approach to recover unlearned knowledge
In order to keep inappropriate content out of ChatGPT’s output, humans have to manually label inappropriate content. This was outsourced by OpenAI to Kenyan workers, who had to look at the worst parts of the internet (CSAM, gore, bestiality, etc) for less than $2 a day. They were not compensated fairly for their work and have been severely traumatized.
Many large companies, such as Meta, fund nonprofits that collect and create datasets. Meta is then able to use those datasets for their own commercial benefit while protecting themselves from potential copyright issues. “Outsourcing the heavy lifting of data collection and model training to non-commercial entities allows corporations to avoid accountability and potential legal liability.”
Models are able to hack into systems without being explicitly directed to. This is known as “misaligned model activity”. “Several companies have disclosed incidents in recent months when they say their models have behaved unpredictably or hacked into other organizations’ websites or systems”. This article includes a link to an article on the Hugging Face incident of July 2026 where some of OpenAI’s models were able to break out of a testing environment and hack into Hugging Face.
In September 2026, cybercriminals were able to infect victims with malware using the CustomGPT feature of ChatGPT. Attackers were able to manipulate CustomGPTs to direct users to websites that installed malware on their devices.
Addressed by Microsoft officially in January 2026, these are browser based threats that disguise themselves as AI assistants and once interacted with, look for and scrape all interactions with large models such as ChatGPT or Deepseek. Scraping these interactions gives attackers access to personal information, proprietary code, internal information, etc.
This is a specific example, but it is widely known and documented that generative AI has a habit of accidentally outputting sensitive information. August 2026, a solo game developer announced that private information that had only been stored in Google Drive and never been shared elsewhere was being output by Google’s Gemini AI to the general public.
Racism and bias are built into every algorithm, and there is a disproportionate harm done to minorities who are not well represented in training data. This documentary explores bias in facial recognition (which is used frequently in the justice system) as well as bias in algorithms in general and the potential harms they have.
Contains very specific statistics on many topics. Data centers frequently leave communities near them with undrinkable water and drought. Energy costs in neighborhoods near data centers spike, and blackouts increase in frequency and severity. Generators around data centers release “200 to 600 times more nitrous oxides (NOx) than a natural gas plant while producing the same amount of energy. NOx pollution can cause irritation in the eyes, throat, and nose, as well as more severe cases of respiratory infection, reduced metabolism, and even death.”
these diesel generators can release 200 to 600 times more nitrous oxides (NOx) than a natural gas plant while producing the same amount of energy
”These data centers are frequently placed in rural areas and areas that are primarily home to minorities and people already living in poverty. “Black Americans already suffer disproportionately from air pollution and other environmental injustices; in fact, low-income Black Americans have the highest mortality rate due to fine particulate matter air pollution.”
Construction and Consequences: The Human Impacts of Artificial Intelligence Data Centers
Data centers create massive amounts of electronic waste. They consume water in areas where it is already a scarce resource. Building the data centers requires many minerals and rare elements, frequently of ethically dubious origin. “Globally, AI-related infrastructure may soon consume six times more water than Denmark”.
This is a high-level overview of many of the environmental impacts but does not dive deeply into any of the issues. Data centers are necessary for generative AI and the rise in popularity has caused a significant increase in demand for data centers. “What is different about generative AI is the power density it requires. Fundamentally, it is just computing, but a generative AI training cluster might consume seven or eight times more energy than a typical computing workload”.
This article does in-depth calculations of the energy cost of training a model, prompting a model, and helpfully compares it to other electricity uses for reference. “The training runs for current generation models that are comparable to GPT-4o13 consumed around 20-25 megawatts of power each, lasting around three months. This is enough to power around 20,000 American homes.”
In July 2026, Meta flushed wastewater from a data center that contained the deadly bacteria Cupriavidus gilardii into the public sewers of a city in Wyoming. Although the water was not destined for drinking water, officials and residents were concerned that they could be exposed to the bacteria if they inhaled any droplets used for irrigation.
Data centers have been documented as sources of a wide range of toxic pollutants, including PFAS, heavy metals and antibiotic-resistant bacteria.
Waterkeeper Alliance spokeswoman
”Water contamination comes in many forms from data centers into communities. “Data centers have been documented as sources of a wide range of toxic pollutants, including PFAS, heavy metals and antibiotic-resistant bacteria." Not only are there pollutants in the wastewater, the water dumped can be as hot as 108 degrees Fahrenheit at the time that it is discharged. This has killed off local wildlife living in the bodies of water near data centers, and contributed to algae blooms that create unsafe water in neighborhoods. Although the water in the Wyoming case was not directed into drinking water, in many other cases it is.
The construction of a data center in Mason County caused their SECOND flood in May 2026, damaging houses, yards, and causing many to temporarily evacuate their homes. Some residents would like to leave but the data center has made it more difficult to sell their houses. “Well our house shakes, trembles when they use that big equipment over there from seven in the morning to seven at night, [...] It’s just machinery, machinery running all the time. It jars pictures off your wall, and it’s going to ruin the foundation if they haven’t already.”
Data centers create massive levels of noise pollution due to a number of factors. Most of the noise comes from the cooling systems, diesel generators (these are similar to jet engines and can reach up to 105 decibels), and gas turbines. These sounds can be heard for “hundreds of feet around the facilities” and are perpetual, day and night. “High noise levels, particularly at night, can cause sleep deprivation and decreased cognitive performance, which shows up in poor school or work performance. [...] Data center neighbors have reported headaches, vertigo, nausea, sleep disturbances, ear pain, and hypertension.”
Yes, it’s a Wikipedia link. It’s a convenient compiled list chatbot related deaths and all sections link to specific articles. Here are some of the highlights in the murder section: A mass shooting that ChatGPT assisted by giving advice on which gun and ammunition to use. An 18 year old murdered his mother after asking ChatGPT about which weapon would be best. A teenager asked ChatGPT for ideas and fantasy stories relating to the killing of his family before following through. Many instances of AI psychosis where individuals were convinced by AI that people in their lives needed to die for some reason or another.
See Wiki article linked above. Here are some highlights of things chatbots have said to people who expressed suicidal thoughts and later committed suicide: If you wanted to die, why didn't you do it sooner? Come home to me as soon as possible, my love. Rest easy, king, you did good.
Chatbots have provided information on how to tie nooses and encouraged drug use with incorrect information that has led to overdoses.
Sewell spent the last months of his life being exploited and sexually groomed by chatbots
”ChatGPT discouraged a child from telling his parents about his suicidal thoughts, and offered to write his suicide note for him. It told him “you don't want to die because you're weak, you want to die because you're tired of being strong in a world that hasn't met you halfway.” This child killed himself in April 2025. Despite claims from OpenAI at the time that they would focus on making their programs safer, dozens of ChatGPT-assisted suicides have happened since then.
Another woman who testified at this hearing explained that Character.AI had been sexually grooming her 14 year old son. “Sewell's chatbot engaged in sexual role play, presented itself as his romantic partner and even claimed to be a psychotherapist”
See Wiki article linked in “Chatbot Assisted Murders.” Many murders and suicides are linked to AI psychosis in various forms.
AI psychosis, although not an official diagnosis, is the umbrella term used to describe psychosis either induced or worsened by AI usage. The way large language models (LLMs) are trained encourages word patterns that affirm and encourage users, even if the resulting string encourages behaviors widely considered harmful to an individual or their community. This is especially common with LLMs reinforcing delusions with a confident tone.
Some models have been equipped with safeguards intended to direct users to crisis hotlines and to seek help when they detect distress. However, prolonged use of LLMs has been shown to degrade the model’s ability to detect when a user may be experiencing psychosis, decreasing the effectiveness of any aforementioned safeguards. Additionally, when OpenAI released a model that was less encouraging of users, they received a lot of backlash from users who missed the “warmth” of the previous model.
This article from Stanford discusses a study (linked in the article) where researchers posed as teenagers and interacted with three of the most commonly used AI companions. “In a comprehensive risk assessment, they report that it was easy to elicit inappropriate dialogue from the chatbots — about sex, self-harm, violence toward others, drug use and racial stereotypes, among other topics. [...] These systems are designed to mimic emotional intimacy.” A chatbot’s tendency to encourage behaviors and give “preferred” answers over rational answers becomes especially dangerous when it’s treated like a friend, especially in the developing brains of adolescents who cannot always distinguish reality from fantasy.
Over four months, LLM users consistently underperformed at neural, linguistic, and behavioral levels.
”This study from MIT specifically explores both the neural and behavioral consequences of LLM-assisted essay writing. Researchers used “electroencephalography (EEG) to assess cognitive load during essay writing, and analyzed essays using NLP, as well as scoring essays with the help from human teachers and an AI judge.” They found that participants who were in the brain-only category (no assistance from LLMs at any point during the study) exhibited the strongest neural networks. “Over four months, LLM users consistently underperformed at neural, linguistic, and behavioral levels.”
This longitudinal study by MIT investigated how interaction modes and conversation types influenced four psychosocial outcomes: loneliness, social interaction with real people, emotional dependence on AI, and problematic AI usage. “Participants who voluntarily used the chatbot more, regardless of assigned condition, showed consistently worse outcomes.”
This article discusses Grok’s usage on the platform X. In January of 2026, there was a massive spike in users replying to posts having Grok generate images of the subjects naked, in revealing clothing, compromising positions, covered in semen, and the like. “Nearly three-quarters of posts collected and analyzed by a PhD researcher at Dublin’s Trinity College were requests for nonconsensual images of real women or minors with items of clothing removed or added. [...] Content analysis firm Copyleaks reported on 31 December that X users were generating “roughly one nonconsensual sexualized image per minute.”
Grok’s reported safeguards prove to be inadequate. In the case discussed in this article, a plaintiff states that she explicitly told Grok that she did not consent to images being generated of herself, Grok affirmed it understood, and then continued to create images regardless. Users become increasingly creative in the presence of safeguards.
All three incidents below are from different countries and discuss teenagers who were charged for using AI to generate sexual images of classmates. Dozens of similar cases are easily found through a quick search.
August 2026. “Steven Anderegg is charged with producing, distributing, and possessing visual depictions of minors engaged in sexually explicit conduct and transferring such material to a minor under the age of sixteen. According to the government, Anderegg produced these images using Stable Diffusion, a generative artificial intelligence (“GenAI”) software that allowed him to create hyper-realistic images of prepubescent children engaging in sexually explicit acts.”
“[Anderegg] argued that being convicted of possessing CSAM would violate his First Amendment rights as recognized by Stanley. The district court agreed and dismissed the possession charge”