Goto

Collaborating Authors

 Media


Textwash -- automated open-source text anonymisation

arXiv.org Artificial Intelligence

With the increasing digitisation of society and human communication, text data are becoming more important for research in the social and behavioural sciences (Gentzkow, Kelly, and Taddy 2019; Salganik 2019). Advances made in natural language processing (NLP) in particular have led to exciting insights derived from text data (e.g., on emotional responses to the pandemic (Kleinberg, Vegt, and Mozes 2020) or on the rhetoric around immigration in political speeches (Card et al. 2022); for an overview, see (Boyd and Schwartz 2021)). Importantly, the use of computational techniques to quantify and analyse text data has triggered a demand, especially for large datasets (often of several tens of thousands of documents) that can be harnessed for machine learning approaches (e.g., (Socher et al. 2013; Lewis et al. 2020)). That status quo of a need for larger datasets and an appetite to use text data for the study of social science phenomena has resulted in a dilemma: many of the important questions require targeted, primary data collection or access to potentially sensitive data. However, such data are hard to obtain, not because they do not exist but because sharing them is constrained by data protection regulations and ethical concerns. One potential consequence is that research activity may be biased toward topics for which suitable data is more readily available rather than those most important. One of the few viable solutions to this dilemma is automated text anonymisation; that is, the large-scale processing of text data so that individuals cannot be identified from the resulting output. Such a method would allow for the flow of sensitive data so that the staggering potential of text data can be exploited for scientific progress. With this paper and the tool it introduces, we seek to enable researchers to work with such sensitive data in a way that protects the privacy of individuals whilst retaining the usefulness of anonymised data for computational text analysis.


The History of AI Rights Research

arXiv.org Artificial Intelligence

This report documents the history of research on AI rights and other moral consideration of artificial entities. It highlights key intellectual influences on this literature as well as research and academic discussion addressing the topic more directly. We find that researchers addressing AI rights have often seemed to be unaware of the work of colleagues whose interests overlap with their own. Academic interest in this topic has grown substantially in recent years; this reflects wider trends in academic research, but it seems that certain influential publications, the gradual, accumulating ubiquity of AI and robotic technology, and relevant news events may all have encouraged increased academic interest in this specific topic. We suggest four levers that, if pulled on in the future, might increase interest further: the adoption of publication strategies similar to those of the most successful previous contributors; increased engagement with adjacent academic fields and debates; the creation of specialized journals, conferences, and research institutions; and more exploration of legal rights for artificial entities.


Label-Efficient Self-Training for Attribute Extraction from Semi-Structured Web Documents

arXiv.org Artificial Intelligence

Extracting structured information from HTML documents is a long-studied problem with a broad range of applications, including knowledge base construction, faceted search, and personalized recommendation. Prior works rely on a few human-labeled web pages from each target website or thousands of human-labeled web pages from some seed websites to train a transferable extraction model that generalizes on unseen target websites. Noisy content, low site-level consistency, and lack of inter-annotator agreement make labeling web pages a time-consuming and expensive ordeal. We develop LEAST -- a Label-Efficient Self-Training method for Semi-Structured Web Documents to overcome these limitations. LEAST utilizes a few human-labeled pages to pseudo-annotate a large number of unlabeled web pages from the target vertical. It trains a transferable web-extraction model on both human-labeled and pseudo-labeled samples using self-training. To mitigate error propagation due to noisy training samples, LEAST re-weights each training sample based on its estimated label accuracy and incorporates it in training. To the best of our knowledge, this is the first work to propose end-to-end training for transferable web extraction models utilizing only a few human-labeled pages. Experiments on a large-scale public dataset show that using less than ten human-labeled pages from each seed website for training, a LEAST-trained model outperforms previous state-of-the-art by more than 26 average F1 points on unseen websites, reducing the number of human-labeled pages to achieve similar performance by more than 10x.


'Sentient' artificial intelligence: Have we reached peak AI hype?

#artificialintelligence

Were you unable to attend Transform 2022? Check out all of the summit sessions in our on-demand library now! Thousands of artificial intelligence experts and machine learning researchers probably thought they were going to have a restful weekend. Then came Google engineer Blake Lemoine, who told the Washington Post on Saturday that he believed LaMDA, Google's conversational AI for generating chatbots based on large language models (LLM), was sentient. Lemoine, who worked for Google's Responsible AI organization until he was placed on paid leave last Monday, and who "became ordained as a mystic Christian priest, and served in the Army before studying the occult," had begun testing LaMDA to see if it used discriminatory or hate speech.


The 10 Best Books About Artificial Intelligence

#artificialintelligence

Long before the technology even existed in the real world, the concept of artificial intelligence has long been a topic of fixation for writers. From cautionary tales and science fiction epics to nonfictional explorations of the implications of AI in our modern world, artificial intelligence seems to be an endlessly fascinating subject of books both big and small. As such, there are all kinds of truly exceptional books about artificial intelligence out there for you to read, enjoy, and maybe even learn a thing or two from. As to be expected, these books about artificial intelligence truly run the gamut. Beyond simply falling under both fiction and nonfiction, artificial intelligence books cover topics ranging from the future to the past, from work to society, from computing to critiques… and all sorts of other topics along the way.


Will Google's AI replace 90% of journalists by 2025?

#artificialintelligence

For over a decade, the newsrooms are facing a daunting challenge to reduce costs to survive in a highly competitive global environment. From smartphone journalism to data journalism to reporting the news on Twitter and TikTok, news outlets without using new technology is unimaginable. With artificial intelligence (AI) entering every avenue of life, both smaller and big newsrooms are vying to employ the new tool and some predictions say by 2025, nearly 90% of news will be written by AI. In content writing, AI is already employed in transcription software, involving recognition and generation of words from an audio file. Currently, several data-based stories are making use of AI, though employing human supervision at the final stage.


The Relonch Camera Leaves Almost Everything Up to AI

#artificialintelligence

Relonch takes the same "we'll develop the photos for you" approach as Kodak's first consumer film camera, editing RAW or digital negative files for the user as a service, though it is one powered by AI and not workers in a darkroom. The colorful leather-wrapped camera itself is simple -- in fact, there is only a shutter button, power button, lens and a viewfinder. Relonch says the camera is designed for everyone who ha never had the time to learn how to use all the controls on their DSLR. Once the user presses that single button, the image is uploaded to the service. The best shots -- not all of them -- are both selected and edited by the AI software, nicknamed Alfred, and sent back to the user at 9 a.m. the next day via the Relonch mobile app.


Camera's artificial intelligence snaps nice photos -- for a price

#artificialintelligence

Still, co-founders Sergey Korzhenrvich and Yuriy Motin on Tuesday are opening an unusual camera shop on University Avenue in downtown Palo Alto. With venture capitalists and entrepreneurs tripping over each other on its sidewalks, the city is as likely a market for their pricey subscriptions as any. Korzhenrvich and Motin are convinced the showroom can help sell customers on the dream of unleashing an artist trapped inside all of us. "If you own a regular camera, you get the hassles," Korzhenrvich said. "In our case, you unleash your talent from inside. You believe there is a great photographer inside you. And you will know that ordinary and mundane is photogenic."


AI Creating 'Art' Is An Ethical And Copyright Nightmare

#artificialintelligence

It's August 2022, and by now you've no doubt read (or more likely seen) something about AI art by now. Whether it's random jokes made for Twitter or paintings that look like they were made by actual human beings, artificial intelligence's ability to create art has exploded onto the scene over the last few months, and while this has been great news for shitposts and fans of tech, it has also raised a number of important questions and concerns. If you haven't read or seen anything about the subject, AI art--or at least as it exists in the state we know it today--is, as Ahmed Elgammal writing in American Scientist so neatly puts it, made when "artists write algorithms not to follow a set of rules, but to'learn' a specific aesthetic by analyzing thousands of images. The algorithm then tries to generate new images in adherence to the aesthetics it has learned." Currently there are a handful of prominent platforms that people are using, with three of the most popular being Midjourney, Dall-E and Stable Diffusion.


The AI startup erasing call center worker accents: is it fighting bias – or perpetuating it?

#artificialintelligence

"Now I have enabled the accent translation," he says. It's the same person, but he sounds completely different: loud and slightly nasal, impossible to distinguish from the accents of my friends in Brooklyn. Only after he had spoken a few more sentences did I notice a hint of the software changing his voice: it rendered the word "technology" with an unnatural cadence and stress on the wrong syllable. Still, it was hard not to be impressed – and disturbed. The man calling me was a product manager from Sanas, a Silicon Valley startup that's building real-time voice-altering technology that aims to help call center workers around the world sound like westerners.