content farm
AI "News" Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian
Puccetti, Giovanni, Rogers, Anna, Alzetta, Chiara, Dell'Orletta, Felice, Esuli, Andrea
Large Language Models (LLMs) are increasingly used as "content farm" models (CFMs), to generate synthetic text that could pass for real news articles. This is already happening even for languages that do not have high-quality monolingual LLMs. We show that fine-tuning Llama (v1), mostly trained on English, on as little as 40K Italian news articles, is sufficient for producing news-like texts that native speakers of Italian struggle to identify as synthetic. We investigate three LLMs and three methods of detecting synthetic texts (log-likelihood, DetectGPT, and supervised classification), finding that they all perform better than human raters, but they are all impractical in the real world (requiring either access to token likelihood information or a large dataset of CFM texts). We also explore the possibility of creating a proxy CFM: an LLM fine-tuned on a similar dataset to one used by the real "content farm". We find that even a small amount of fine-tuning data suffices for creating a successful detector, but we need to know which base LLM is used, which is a major challenge. Our results suggest that there are currently no practical methods for detecting synthetic news-like texts 'in the wild', while generating them is too easy. We highlight the urgency of more NLP research on this problem.
Is computational creativity flourishing on the dead internet?
T erence Broad Creative Computing Institute University of the Arts London United Kingdom t.broad@arts.ac.uk Abstract The dead internet theory is a conspiracy theory that states that all interactions and posts on social media are no longer being made by real people, but rather by autonomous bots. While the theory is obviously not true, an increasing amount of posts on social media have been made by bots optimised to gain followers and drive engagement on social media platforms. This paper looks at the recent phenomenon of these bots, analysing their behaviour through the lens of computational creativity to investigate the question: is computational creativity flourishing on the dead internet? Introduction The dead internet theory is a conspiracy theory that emerged in the late 2010's or early 2020's that states that large parts of the internet, in particular on social media are no longer occupied by humans and human generated content, but rather posts by AI-driven bots that are designed to control or influence human behaviour (IlluminatiPirate 2021). Whist the theory emerges from the fringes of the internet, stemming in conspiratorial thinking as a way of explaining broad-based changes to society from nefarious actors, many commentators have observed that there is a grain of truth to the theory (Tiffany 2021).
Amazon to crack down on self-publishers using AI-generated content
The'America's Got Talent' judge told Fox News Digital why he doesn't like AI technology in songwriting. Amazon will require publishers on Kindle to disclose when any of their content is generated by artificial intelligence after complaints forced the company to take action. "We require you to inform us of AI-generated content (text, images or translations) when you publish a new book or make edits to and republish an existing book through KDP (Kindle Direct Publishing). AI-generated images include cover and interior images and artwork," Amazon said of the updated guidelines, according to a report in Cyber News. The update comes after the company faced complaints from users that some works being sold under the names of human writers contained content that was either fully or partially generated by AI, according to the report.
Redditors troll an AI content farm into covering a fake 'WoW' feature
Some redditors seem very excited about a new World of Warcraft feature called Glorbo, which some believe will "make a huge impact on the game." Their palpable enthusiasm for Glorbo caught the attention of a blog named The Portal, which publishes "gaming content powered by Z League," an app that aims to bring gamers together. The Portal appears to be using AI to scrape Reddit posts and turn them into content. Redditor u/kaefer_kriegerin noticed that The Portal was seemingly turning discussions from some gaming subreddits into blog posts. They decided to try and trick the content farm into covering a fake WoW feature. The ruse was a success.
Chatbot 'journalists' found running almost 50 AI-generated content farms
Chatbots pretending to be journalists have been discovered running almost 50 AI-generated "content farms" so far, according to an investigation by the anti-misinformation outfit NewsGuard. The websites churn out content relating to politics, health, environment, finance and technology at a "high volume", the researchers found, to provide rapid turnover of material to saturate with adverts for profit. "Some publish hundreds of articles a day," Newsguard's McKenzie Sadeghi and Lorenzo Arvanitis said. In total, 49 sites in seven languages โ English, Chinese, Czech, French, Portuguese, Tagalog and Thai โ were identified as being "entirely or mostly" generated by AI language models. Almost half the sites had no obvious record of ownership or control, and only four were able to be contacted.
SEO trends and Google changes to expect in 2018
We're already over a week into 2018, and the start of a new year is a great time to check in and see where we stand as an industry -- and how things might change this year. Back in 2010, Google was getting beaten up in the media for the increasing amount of "content farm" clutter in the search results. Soon after that, in February 2011, the Google Panda update was released, which specifically targeted spammy and low-quality content. Why do I bring this up today? Because the media has been hammering Google for promoting fake news for the past year and a half -- a problem so extensive that search industry expert Danny Sullivan has referred to it as "Google's biggest-ever search quality crisis."