Government
Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content
O'Neill, Charles, Miller, Jack, Ciuca, Ioana, Ting, Yuan-Sen, Bui, Thang
In this paper, we tackle the emerging challenge of unintended harmful content generation in Large Language Models (LLMs) with a novel dual-stage optimisation technique using adversarial fine-tuning. Our two-pronged approach employs an adversarial model, fine-tuned to generate potentially harmful prompts, and a judge model, iteratively optimised to discern these prompts. In this adversarial cycle, the two models seek to outperform each other in the prompting phase, generating a dataset of rich examples which are then used for fine-tuning. This iterative application of prompting and fine-tuning allows continuous refinement and improved performance. The performance of our approach is evaluated through classification accuracy on a dataset consisting of problematic prompts not detected by GPT-4, as well as a selection of contentious but unproblematic prompts. We show considerable increase in classification accuracy of the judge model on this challenging dataset as it undergoes the optimisation process. Furthermore, we show that a rudimentary model \texttt{ada} can achieve 13\% higher accuracy on the hold-out test set than GPT-4 after only a few rounds of this process, and that this fine-tuning improves performance in parallel tasks such as toxic comment identification.
Counterfactual Reasoning for Bias Evaluation and Detection in a Fairness under Unawareness setting
Cornacchia, Giandomenico, Anelli, Vito Walter, Narducci, Fedelucio, Ragone, Azzurra, Di Sciascio, Eugenio
Current AI regulations require discarding sensitive features (e.g., gender, race, religion) in the algorithm's decision-making process to prevent unfair outcomes. However, even without sensitive features in the training set, algorithms can persist in discrimination. Indeed, when sensitive features are omitted (fairness under unawareness), they could be inferred through non-linear relations with the so called proxy features. In this work, we propose a way to reveal the potential hidden bias of a machine learning model that can persist even when sensitive features are discarded. This study shows that it is possible to unveil whether the black-box predictor is still biased by exploiting counterfactual reasoning. In detail, when the predictor provides a negative classification outcome, our approach first builds counterfactual examples for a discriminated user category to obtain a positive outcome. Then, the same counterfactual samples feed an external classifier (that targets a sensitive feature) that reveals whether the modifications to the user characteristics needed for a positive outcome moved the individual to the non-discriminated group. When this occurs, it could be a warning sign for discriminatory behavior in the decision process. Furthermore, we leverage the deviation of counterfactuals from the original sample to determine which features are proxies of specific sensitive information. Our experiments show that, even if the model is trained without sensitive features, it often suffers discriminatory biases.
The Heated Debate Over Who Should Control Access to AI
In May, the CEOs of three of the most prominent AI labs--OpenAI, Google DeepMind, and Anthropic--signed a statement that warned AI could be as risky to humanity as pandemics and nuclear war. To prevent disaster, many AI companies and researchers are arguing for restrictions on who can access the most powerful AI models and who can develop them in the first place. They worry that bad actors could use AI models to create large amounts of disinformation that could alter the outcomes of elections, and that in the future, more powerful AI models could help launch cyberattacks or create bioweapons. But not all AI companies agree. On Thursday, Meta released Code Llama, a family of AI models built on top of Llama 2, Meta's flagship large language model, with extra training to make them particularly useful for coding tasks.
TikToker sounds alarm on this scary online trend that turns your children into bait for predators
A TikToker warned of a growing trend involving child predators who use artificial intelligence to turn photos and videos of kids into explicit content. Posting imagery of children on social media can invite "digital kidnappers" to steal their likeness and use them in exploitative AI-generated videos, Alex Hoffman said in a viral TikTok video. "Digital kidnapping is when somebody steals the photos of your minor from the internet, usually a social media platform, and either pretends to be the child or pretends to be the child's parents," she said. "Oftentimes digital kidnappers will take normal photos of a child on the internet and alter them to look explicit or show the child doing something inappropriate." "Digital kidnappers can also take photos of a child and make them into an inappropriate video using AI materials," said Hoffman, a law student who has worked with the government investigating online sex crimes against children.
Chris Christie calls out Vivek Ramaswamy for GOP primary debate performance: Uses 'ChatGPT phrases'
Former New Jersey Gov. Chris Christie tore into GOP presidential candidate Vivek Ramaswamy one day after the first primary debate in Milwaukee, arguing the entrepreneur's answers showed he has "absolutely no idea what he's talking about." Christie and Ramaswamy sparred over several issues during the two-hour debate from the United States' role in funding the war in Ukraine to supporting former President Donald Trump if he's convicted. Trump praised Ramaswamy's debate performance on his social media site Truth Social. Ramaswamy also praised Trump on stage as the "best president of the 21st century." "Well, I'm stunned that as I was talking about Donald Trump and all the ways that he's let down our party and our country, that he [Trump] didn't mention me as a winner of the debate last night," Christie said Thursday on "Your World."
We need to avoid a 'ready, fire, aim!' approach to AI regulation
Sam Altman, the CEO of artificial intelligence lab OpenAI, told a Senate panel he welcomes federal regulation on the technology "to mitigate" its risks. The panic to regulate artificial intelligence (AI) came almost immediately after last fall's release of ChatGPT popularized the technology with the public. Some industry insiders themselves called for a pause on development, highlighting that expertise in a field doesn't translate into proficiency in the perils of regulation. That appeal was followed by a White House AI Bill of Rights and an educational effort by Senate Majority Leader Chuck Schumer, D-N.Y. Fears about AI include job displacement, data security and privacy, misinformation, autonomous defense systems mistakes, discrimination and bias, and an existential threat to humanity itself. It's imperative to prove actual market failure before regulating and to make sure the costs of doing so don't outweigh the benefits.
Russia says destroyed 42 Ukraine-launched drones over Crimea
Russia's defence ministry has said its air defence forces destroyed a large-scale Ukrainian-launched drone attack on the Crimean Peninsula, which Moscow annexed from Ukraine in 2014. Crimea has been targeted by Kyiv since Moscow launched its full-scale invasion of Ukraine in February 2022, but has come under more intense, increased attacks in recent weeks. The Russian Ministry of Defence said early on Friday its forces shot down nine drones, while 33 others "were suppressed by electronic warfare and crashed without reaching the target". It did not elaborate on whether there had been any damage or casualties. It added that it had also shot down a Ukraine-launched missile over the Kaluga region, which borders the Moscow region.
New York Times, CNN and Australia's ABC block OpenAI's GPTBot web crawler from accessing content
News outlets including the New York Times, CNN, Reuters and the Australian Broadcasting Corporation (ABC) have blocked a tool from OpenAI, limiting the company's ability to continue accessing their content. OpenAI is behind one of the best known artificial intelligence chatbots, ChatGPT. Its web crawler โ known as GPTBot โ may scan webpages to help improve its AI models. The Verge was first to report the New York Times had blocked GPTBot on its website. The Guardian subsequently found that other major news websites, including CNN, Reuters, the Chicago Tribune, the ABC and Australian Community Media (ACM) brands such as the Canberra Times and the Newcastle Herald, appear to have also disallowed the web crawler.
New York Times, CNN and ABC block OpenAI's GPTBot web crawler from accessing content
News outlets including the New York Times, CNN, Reuters and the Australian Broadcasting Corporation (ABC) have blocked a tool from OpenAI, limiting the company's ability to continue accessing their content. OpenAI is behind one of the best known artificial intelligence chatbots, ChatGPT. Its web crawler โ known as GPTBot โ may scan webpages to help improve its AI models. The Verge was first to report the New York Times had blocked GPTBot on its website. The Guardian subsequently found that other major news websites, including CNN, Reuters, the Chicago Tribune, the ABC and Australian Community Media (ACM) brands such as the Canberra Times and the Newcastle Herald, appear to have also disallowed the web crawler.
GeoExplainer: A Visual Analytics Framework for Spatial Modeling Contextualization and Report Generation
Lei, Fan, Ma, Yuxin, Fotheringham, Stewart, Mack, Elizabeth, Li, Ziqi, Sachdeva, Mehak, Bardin, Sarah, Maciejewski, Ross
Geographic regression models of various descriptions are often applied to identify patterns and anomalies in the determinants of spatially distributed observations. These types of analyses focus on answering why questions about underlying spatial phenomena, e.g., why is crime higher in this locale, why do children in one school district outperform those in another, etc.? Answers to these questions require explanations of the model structure, the choice of parameters, and contextualization of the findings with respect to their geographic context. This is particularly true for local forms of regression models which are focused on the role of locational context in determining human behavior. In this paper, we present GeoExplainer, a visual analytics framework designed to support analysts in creating explanative documentation that summarizes and contextualizes their spatial analyses. As analysts create their spatial models, our framework flags potential issues with model parameter selections, utilizes template-based text generation to summarize model outputs, and links with external knowledge repositories to provide annotations that help to explain the model results. As analysts explore the model results, all visualizations and annotations can be captured in an interactive report generation widget. We demonstrate our framework using a case study modeling the determinants of voting in the 2016 US Presidential Election.