Media
Reliably detecting AI-generated text is mathematically impossible
Determining whether text has come from artificial intelligence models like ChatGPT might be impossible to do reliably, according to a new mathematical proof. The ease with which AI models generate text that seems as if it were written by a human has led to issues such as cheating on essays and exams, and mass disinformation campaigns.
Debt Rattle March 30 2023 - The Automatic Earth
Carried out the 2014 coup d'état in Ukraine The US is a state sponsor of terrorism https://t.co/CKykvqUa5U To be discussed tonight pic.twitter.com/kTf0ehwhbG This is one of the most disturbing videos I have ever seen. It confirms that the TGA knew back in Jan 2021 that the lipid nanoparticles (and the mRNA) didn't stay in the inject site, but spread throughout the entire body including the brain, the liver and female ovaries. Today, Democrats defeated my amendment to require Senate ratification for any pandemic agreement with the World Health Organization. Now we know Democrats are willing to relinquish U.S. sovereignty to a global entity. Jim lays it out very well. Renowned author and journalist James Howard Kunstler (JHK) has been complaining and pointing out that the American public is told one lie after another by the Lying Legacy Media (LLM), the government and the medical community. This kind of lying, according to JHK, is pure treason by all parties, from the 600 million CV19 bioweapon/vax injections, to the crumbling banking system, to the war in Ukraine. Let's start with the genocide of the CV19vax.
The Complete Visual Guide to Machine Learning and Data Science - CouponED
In Part 1 we'll introduce the machine learning workflow and common techniques for cleaning and preparing raw data for analysis. We'll explore univariate analysis with frequency tables, histograms, kernel densities, and profiling metrics, then dive into multivariate profiling tools like heat maps, violin and box plots, scatter plots, and correlation: Variable types, empty values, range and count calculations, left/right censoring, etc. Histograms, frequency tables, mean, median, mode, variance, skewness, etc. Throughout the course, we'll introduce real-world scenarios to solidify key concepts and simulate actual data science and business intelligence cases. You'll use profiling metrics to clean up product inventory data for a local grocery, explore Olympic athlete demographics with histograms and kernel densities, visualize traffic accident frequency with heat maps, and more. In Part 2 we'll introduce the supervised learning landscape, review the classification workflow, and address key topics like dependent vs. independent variables, feature engineering, data splitting and overfitting.
What can Google's AI-powered Bard do? We tested it for you
To use, or not to use, Bard? That is the Shakespearean question an Associated Press reporter sought to answer while testing out Google's artificially intelligent chatbot. The recently rolled-out bot dubbed Bard is the internet search giant's answer to the ChatGPT tool that Microsoft has been melding into its Bing search engine and other software. During several hours of interaction, the AP learned Bard is quite forthcoming about its unreliability and other shortcomings, including its potential for mischief in next year's U.S. presidential election. Even as it occasionally warned of the problems it could unleash, Bard repeatedly emphasized its belief that it will blossom into a force for good.
AI in marketing -- The amazing potential and real limitations
There's so much buzz about AI in marketing I need a bee keeper's suit just to keep the bullshit off of me. The marketing world is swarming with articles, opinions, podcasts and videos about how AI's going to change world completely. Most of the publicity is positive, touting the time savings and efficiency that AI tools will provide. But there are also plenty of Chicken Littles who are saying I'm bound to lose my job any day now. That kind of fear is a familiar refrain for those of us who know the history of marketing. Way back in the 50s when television was widely adopted, everyone said radio was dead. They said it again in 1981 when MTV came out… "Who would want to just listen to music when you can watch music videos."
Using artificial intelligence and archival news articles, this teen found that Black homicide victims were less humanized in news coverage
Using artificial intelligence and archival news articles, a teenager in Northern Virginia created a program to measure media biases – and in researching older news articles, she found that Black homicide victims were less likely to be humanized in news coverage. Emily Ocasio, an 18-year-old from Falls Church, Virginia, created an AI program that analyzed FBI homicide records between 1976 and 1984 and their corresponding coverage published in The Boston Globe to determine whether victims were presented in a humanizing or impersonal way. After analyzing 5,042 entries, the results showed that Black men under the age of 18 were 30% less likely to receive humanizing coverage than their White counterparts, Ocasio told CNN. Black women were 23% less likely to be humanized in news stories, Ocasio added. A news article was considered humanizing when it mentioned additional information about the victim and presented them "as a person, not just a statistic," Ocasio said in her project presentation.
Yes but.. Can ChatGPT Identify Entities in Historical Documents?
González-Gallardo, Carlos-Emiliano, Boros, Emanuela, Girdhar, Nancy, Hamdi, Ahmed, Moreno, Jose G., Doucet, Antoine
Large language models (LLMs) have been leveraged for several years now, obtaining state-of-the-art performance in recognizing entities from modern documents. For the last few months, the conversational agent ChatGPT has "prompted" a lot of interest in the scientific community and public due to its capacity of generating plausible-sounding answers. In this paper, we explore this ability by probing it in the named entity recognition and classification (NERC) task in primary sources (e.g., historical newspapers and classical commentaries) in a zero-shot manner and by comparing it with state-of-the-art LM-based systems. Our findings indicate several shortcomings in identifying entities in historical text that range from the consistency of entity annotation guidelines, entity complexity, and code-switching, to the specificity of prompting. Moreover, as expected, the inaccessibility of historical archives to the public (and thus on the Internet) also impacts its performance.
Detecting and Grounding Important Characters in Visual Stories
Characters are essential to the plot of any story. Establishing the characters before writing a story can improve the clarity of the plot and the overall flow of the narrative. However, previous work on visual storytelling tends to focus on detecting objects in images and discovering relationships between them. In this approach, characters are not distinguished from other objects when they are fed into the generation pipeline. The result is a coherent sequence of events rather than a character-centric story. In order to address this limitation, we introduce the VIST-Character dataset, which provides rich character-centric annotations, including visual and textual co-reference chains and importance ratings for characters. Based on this dataset, we propose two new tasks: important character detection and character grounding in visual stories. For both tasks, we develop simple, unsupervised models based on distributional similarity and pre-trained vision-and-language models. Our new dataset, together with these models, can serve as the foundation for subsequent work on analysing and generating stories from a character-centric perspective.
The Music Annotation Pattern
de Berardinis, Jacopo, Meroño-Peñuela, Albert, Poltronieri, Andrea, Presutti, Valentina
The annotation of music content is a complex process to represent due to its inherent multifaceted, subjectivity, and interdisciplinary nature. Numerous systems and conventions for annotating music have been developed as independent standards over the past decades. Little has been done to make them interoperable, which jeopardises cross-corpora studies as it requires users to familiarise with a multitude of conventions. Most of these systems lack the semantic expressiveness needed to represent the complexity of the musical language and cannot model multi-modal annotations originating from audio and symbolic sources. In this article, we introduce the Music Annotation Pattern, an Ontology Design Pattern (ODP) to homogenise different annotation systems and to represent several types of musical objects (e.g. chords, patterns, structures). This ODP preserves the semantics of the object's content at different levels and temporal granularity. Moreover, our ODP accounts for multi-modality upfront, to describe annotations derived from different sources, and it is the first to enable the integration of music datasets at a large scale.
On the Complexity of Finding Set Repairs for Data-Graphs
Abriola, Sergio (Conicet UBA) | Martínez, María Vanina (Conicet UBA) | Pardal, Nina (Conicet UBA) | Cifuentes, Santiago (a:1:{s:5:"en_US";s:8:"FCEN UBA";}) | Pin Baque, Edwin (FCEN UBA)
In the deeply interconnected world we live in, pieces of information link domains all around us. As graph databases embrace effectively relationships among data and allow processing and querying these connections efficiently, they are rapidly becoming a popular platform for storage that supports a wide range of domains and applications. As in the relational case, it is expected that data preserves a set of integrity constraints that define the semantic structure of the world it represents. When a database does not satisfy its integrity constraints, a possible approach is to search for a ‘similar’ database that does satisfy the constraints, also known as a repair. In this work, we study the problem of computing subset and superset repairs for graph databases with data values using a notion of consistency based on having a set of Reg-GXPath expressions as integrity constraints. We show that for positive fragments of Reg-GXPath these problems admit a polynomial-time algorithm, while the full expressive power of the language renders them intractable.