Media
Improving Fake News Detection of Influential Domain via Domain- and Instance-Level Transfer
Nan, Qiong, Wang, Danding, Zhu, Yongchun, Sheng, Qiang, Shi, Yuhui, Cao, Juan, Li, Jintao
Both real and fake news in various domains, such as politics, health, and entertainment are spread via online social media every day, necessitating fake news detection for multiple domains. Among them, fake news in specific domains like politics and health has more serious potential negative impacts on the real world (e.g., the infodemic led by COVID-19 misinformation). Previous studies focus on multi-domain fake news detection, by equally mining and modeling the correlation between domains. However, these multi-domain methods suffer from a seesaw problem: the performance of some domains is often improved at the cost of hurting the performance of other domains, which could lead to an unsatisfying performance in specific domains. To address this issue, we propose a Domain- and Instance-level Transfer Framework for Fake News Detection (DITFEND), which could improve the performance of specific target domains. To transfer coarse-grained domain-level knowledge, we train a general model with data of all domains from the meta-learning perspective. To transfer fine-grained instance-level knowledge and adapt the general model to a target domain, we train a language model on the target domain to evaluate the transferability of each data instance in source domains and re-weigh each instance's contribution. Offline experiments on two datasets demonstrate the effectiveness of DITFEND. Online experiments show that DITFEND brings additional improvements over the base models in a real-world scenario.
Noise-Robust De-Duplication at Scale
Silcock, Emily, D'Amico-Wong, Luca, Yang, Jinglin, Dell, Melissa
Identifying near duplicates within large, noisy text corpora has a myriad of applications that range from de-duplicating training datasets, reducing privacy risk, and evaluating test set leakage, to identifying reproduced news articles and literature within large corpora. Across these diverse applications, the overwhelming majority of work relies on N-grams. Limited efforts have been made to evaluate how well N-gram methods perform, in part because it is unclear how one could create an unbiased evaluation dataset for a massive corpus. This study uses the unique timeliness of historical news wires to create a 27,210 document dataset, with 122,876 positive duplicate pairs, for studying noise-robust de-duplication. The time-sensitivity of news makes comprehensive hand labelling feasible - despite the massive overall size of the corpus - as duplicates occur within a narrow date range. The study then develops and evaluates a range of de-duplication methods: hashing and N-gram overlap (which predominate in the literature), a contrastively trained bi-encoder, and a re-rank style approach combining a bi- and cross-encoder. The neural approaches significantly outperform hashing and N-gram overlap. We show that the bi-encoder scales well, de-duplicating a 10 million article corpus on a single GPU card in a matter of hours. The public release of our NEWS-COPY de-duplication dataset will facilitate further research and applications.
TVStoryGen: A Dataset for Generating Stories with Character Descriptions
We introduce TVStoryGen, a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved. Unlike other story generation datasets, TVStoryGen contains stories that are authored by professional screen-writers and that feature complex interactions among multiple characters. Generating stories in TVStoryGen requires drawing relevant information from the lengthy provided documents about characters based on the brief summary. In addition, we propose to train reverse models on our dataset for evaluating the faithfulness of generated stories. We create TVStoryGen from fan-contributed websites, which allows us to collect 26k episode recaps with 1868.7 tokens on average. Empirically, we take a hierarchical story generation approach and find that the neural model that uses oracle content selectors for character descriptions demonstrates the best performance on automatic metrics, showing the potential of our dataset to inspire future research on story generation with constraints. Qualitative analysis shows that the best-performing model sometimes generates content that is unfaithful to the short summaries, suggesting promising directions for future work.
Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation
Xu, Jin, Liu, Xiaojiang, Yan, Jianhao, Cai, Deng, Li, Huayang, Li, Jian
While large-scale neural language models, such as GPT2 and BART, have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (\textit{e.g.}, greedy search). This phenomenon is counter-intuitive since there are few consecutive sentence-level repetitions in human corpora (e.g., 0.02\% in Wikitext-103). To investigate the underlying reasons for generating consecutive sentence-level repetitions, we study the relationship between the probabilities of the repetitive tokens and their previous repetitions in the context. Through our quantitative experiments, we find that 1) Language models have a preference to repeat the previous sentence; 2) The sentence-level repetitions have a \textit{self-reinforcement effect}: the more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence; 3) The sentences with higher initial probabilities usually have a stronger self-reinforcement effect. Motivated by our findings, we propose a simple and effective training method \textbf{DITTO} (Pseu\underline{D}o-Repet\underline{IT}ion Penaliza\underline{T}i\underline{O}n), where the model learns to penalize probabilities of sentence-level repetitions from pseudo repetitive data. Although our method is motivated by mitigating repetitions, experiments show that DITTO not only mitigates the repetition issue without sacrificing perplexity, but also achieves better generation quality. Extensive experiments on open-ended text generation (Wikitext-103) and text summarization (CNN/DailyMail) demonstrate the generality and effectiveness of our method.
4 Types of Machine Learning and Explained
Machine learning is a significant area of technology, as you are aware. The majority of technology today is evolving toward artificial intelligence (which are created on Machine Learning). The majority of businesses are more focused on AI. Deep Learning is a different, more sophisticated sort of learning. You need to grasp what machine learning and AI are before diving further into these four forms of machine learning.
AI music generators could be a boon for artists -- but also problematic
It was only five years ago that electronic punk band YACHT entered the recording studio with a daunting task: They would train an AI on 14 years of their music, then synthesize the results into the album "Chain Tripping." "I'm not interested in being a reactionary," YACHT member and tech writer Claire L. Evans said in a documentary about the album. "I don't want to return to my roots and play acoustic guitar because I'm so freaked out about the coming robot apocalypse, but I also don't want to jump into the trenches and welcome our new robot overlords either." But our new robot overlords are making a whole lot of progress in the space of AI music generation. Even though the Grammy-nominated "Chain Tripping" was released in 2019, the technology behind it is already becoming outdated.
'Chat' with Musk or Trump on AI chatbot
A new chatbot start-up from two top artificial intelligence talents lets anyone strike up a conversation with impersonations of Donald Trump, Elon Musk, Albert Einstein and Sherlock Holmes. Registered users type in messages and get responses. They can also create a chatbot of their own on Character.ai, "There were reports of possible voter fraud and I wanted an investigation," the Trump bot said. The start-up's two founders helped create Google's artificial intelligence project LaMDA, which Google keeps closely guarded while it develops safeguards against social risks.
'It's a living organism, a crazy cacophony of life': Scott A Woodward's best phone picture
Nicknamed the Monster Building, the residential complex in Hong Kong's Quarry Bay is actually made up of five imposing tower blocks. In 2018, Canadian photographer Scott A Woodward had set up camp in the shadow of one, the Yick Cheong building, to shoot an ad campaign for Foot Locker. "It's a heavy, teeming, living organism; a crazy cacophony of life and colour," he says. "There are 10,000 people living there, and people travel from all over to see it." The team was large and busy, vying for space in the courtyard with tourists and Instagrammers drawn to the building after it featured in the films Transformers: Age of Extinction and Ghost in the Shell.
Deepfake Bruce Willis may be the next Hollywood star, and he's OK with that [Updated]
According to the BBC and the Hollywood Reporter, a representative for Bruce Willis said, "Please know that Bruce has no partnership or agreement with this Deepcake company." The original Telegraph report we cited appears to be in error, and it's unclear whether Deepcake ever had the permission to use Willis' likeness beyond a 2021 Russian cell phone commercial. We have published a new piece with more details. Bruce Willis has sold the "digital twin" rights to his likeness for commercial video production use, according to a report by The Telegraph. This move allows the Hollywood actor to digitally appear in future commercials and possibly even films, and he has already appeared in a Russian commercial using the technology.
Who's Who in Data Science and Machine Learning? - Onalytica
Data Science combines statistical and computational skills together with Machine Learning for data-driven problem solving. This rapidly growing area includes large-scale data analysis, DevOps and deep learning, and has applications in many tech and finance related areas, amongst others. Data Science and Machine Learning distinguishes itself from other computer guided decision methods by creating prediction algorithms using data. Well known examples of these applications include spam detectors, movie and music recommendation systems as well as speech recognition tools. Data Science and Machine Learning platforms provide a variety of basic building blocks that aid in creating different types of data science solutions.