Goto

Collaborating Authors

 Government


On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse

arXiv.org Machine Learning

Specifying all desirable properties of a language model is challenging, but certain requirements seem essential. Given samples from an unknown language, the trained model should produce valid strings not seen in training and be expressive enough to capture the language's full richness. Otherwise, outputting invalid strings constitutes "hallucination," and failing to capture the full range leads to "mode collapse." We ask if a language model can meet both requirements. We investigate this within a statistical language generation setting building on Gold and Angluin. Here, the model receives random samples from a distribution over an unknown language K, which belongs to a possibly infinite collection of languages. The goal is to generate unseen strings from K. We say the model generates from K with consistency and breadth if, as training size increases, its output converges to all unseen strings in K. Kleinberg and Mullainathan [KM24] asked if consistency and breadth in language generation are possible. We answer this negatively: for a large class of language models, including next-token prediction models, this is impossible for most collections of candidate languages. This contrasts with [KM24]'s result, showing consistent generation without breadth is possible for any countable collection of languages. Our finding highlights that generation with breadth fundamentally differs from generation without breadth. As a byproduct, we establish near-tight bounds on the number of samples needed for generation with or without breadth. Finally, our results offer hope: consistent generation with breadth is achievable for any countable collection of languages when negative examples (strings outside K) are available alongside positive ones. This suggests that post-training feedback, which encodes negative examples, can be crucial in reducing hallucinations while limiting mode collapse.


How to Get Started on Bluesky

WIRED

The social media app Bluesky just reached the top of the free download charts for Apple's app store in the United States, making it--for the moment, at least--more popular than Meta's Threads and OpenAI's ChatGPT. The decentralized social media platform has received a fresh influx of users over the last week as more X users sour on Elon Musk's political ambitions and abandon his social media platform for an alternative. Over a million new users have joined Bluesky since the 2024 US presidential election on November 5, which was shaped by Musk's influence. First launched in 2019 as a project within Twitter, Bluesky gained independence from the company before Musk's acquisition and its subsequent name change. Bluesky also captured the attention of some ex-tweeters back in 2023, when new users were only able to sign up through an invite system.


Fox News AI Newsletter: AI developers discover 'Donald Trump neuron', expert says

FOX News

Kurt'CyberGuy' Knutsson on President-elect Trump's plan to deregulate cryptocurrency and A.I. in his second administration. 'DONALD TRUMP NEURON': Artificial intelligence recognizes images and the name of President-elect Donald Trump so much that the phenomenon is referred to as a "Donald Trump neuron," expert Chris Olah says. MUSK PETITION: An artificial intelligence (AI) advocacy group is urging President-elect Trump to make billionaire entrepreneur Elon Musk a special adviser to the White House focused on AI. INDIA - 2024/05/17: In this photo illustration, the OpenAI logo is seen displayed on a mobile phone screen with ChatGPT logo in the background. HELP FROM SILICON VALLEY: OpenAI has assembled a "blueprint" for artificial intelligence infrastructure that the company hopes will be considered by the incoming Trump administration and Congress – suggesting that the plan will help the United States maintain its lead in the field over competitors like China.


OpenAI touts AI infrastructure 'blueprint' to outcompete China, bolster economy under incoming Trump admin

FOX News

Kurt'CyberGuy' Knutsson on President-elect Trump's plan to deregulate cryptocurrency and A.I. in his second administration. OpenAI has assembled a "blueprint" for artificial intelligence (AI) infrastructure that the company hopes will be considered by the incoming Trump administration and Congress – suggesting that the plan will help the United States maintain its lead in the field over competitors like China. The company's Vice President of Global Affairs, Chris Lehane, announced the "Infrastructure Blueprint for the U.S." on Wednesday during an event hosted by the Center for Strategic and International Studies (CSIS). The company says AI's potential presents an "unmissable opportunity to revitalize the American Dream and reindustrialize the US." "Investments to extend the current U.S. lead in AI will yield tens of thousands of skilled-trade and other jobs, growth in productivity and GDP; a modernized grid including power generated by nuclear energy; a state-of-the-art network of semiconductor manufacturing facilities; and a new generation of AI-powered businesses and entrepreneurship," OpenAI claims. In this photo illustration, the OpenAI logo is seen displayed on a mobile phone screen with ChatGPT logo in the background.


The words and phrases you should NEVER Google or your computer could get hacked

Daily Mail - Science & tech

Searching on Google might seem like one of the safest things to do online. But cybersecurity experts warn that there are some searches which could put you at serious risk of being hacked. Last week, it was revealed that cybercriminals had hijacked the Google results for'Are Bengal cats legal in Australia?' to infect cat-lovers' computers. Now, experts have revealed the seven other common words and phrases you should never Google. Using a technique called'SEO poisoning' criminals exploit Google's search results to lure unsuspecting victims into websites they control.


Houthis launch missile, drone attacks on US warships off Yemen's coast

Al Jazeera

US warships came under sustained missile and drone attack from Houthi fighters as they sailed off the coast of Yemen, the Pentagon has confirmed, with the armed group claiming it attacked the US aircraft carrier Abraham Lincoln and two US destroyers. Pentagon spokesperson Air Force Major General Patrick Ryder said on Tuesday that the United States military's Central Command (CENTCOM) forces "successfully repelled multiple Iranian backed Houthi attacks during a transit of the Bab al-Mandeb strait", which connects the Red Sea to the Gulf of Aden. Ryder told reporters at a news conference that two US-guided missile destroyers – the USS Stockdale and USS Spruance – were attacked by at least eight one-way attack drones, five antiship ballistic missiles and three antiship cruise missiles. All the Houthi drones and missiles "were successfully engaged and defeated", and neither of the US Navy ships were damaged or personnel hurt, he said. Ryder added that he was not aware of any attacks against the aircraft carrier USS Abraham Lincoln.



PICZL: Image-based Photometric Redshifts for AGN

arXiv.org Machine Learning

Computing photo-z for AGN is challenging, primarily due to the interplay of relative emissions associated with the SMBH and its host galaxy. SED fitting methods, effective in pencil-beam surveys, face limitations in all-sky surveys with fewer bands available, lacking the ability to capture the AGN contribution to the SED accurately. This limitation affects the many 10s of millions of AGN clearly singled out and identified by SRG/eROSITA. Our goal is to significantly enhance photometric redshift performance for AGN in all-sky surveys while avoiding the need to merge multiple data sets. Instead, we employ readily available data products from the 10th Data Release of the Imaging Legacy Survey for DESI, covering > 20,000 deg$^{2}$ with deep images and catalog-based photometry in the grizW1-W4 bands. We introduce PICZL, a machine-learning algorithm leveraging an ensemble of CNNs. Utilizing a cross-channel approach, the algorithm integrates distinct SED features from images with those obtained from catalog-level data. Full probability distributions are achieved via the integration of Gaussian mixture models. On a validation sample of 8098 AGN, PICZL achieves a variance $\sigma_{\textrm{NMAD}}$ of 4.5% with an outlier fraction $\eta$ of 5.6%, outperforming previous attempts to compute accurate photo-z for AGN using ML. We highlight that the model's performance depends on many variables, predominantly the depth of the data. A thorough evaluation of these dependencies is presented in the paper. Our streamlined methodology maintains consistent performance across the entire survey area when accounting for differing data quality. The same approach can be adopted for future deep photometric surveys such as LSST and Euclid, showcasing its potential for wide-scale realisation. With this paper, we release updated photo-z (including errors) for the XMM-SERVS W-CDF-S, ELAIS-S1 and LSS fields.


Optimisation Strategies for Ensuring Fairness in Machine Learning: With and Without Demographics

arXiv.org Artificial Intelligence

Ensuring fairness has emerged as one of the primary concerns in AI and its related algorithms. Over time, the field of machine learning fairness has evolved to address these issues. This paper provides an extensive overview of this field and introduces two formal frameworks to tackle open questions in machine learning fairness. In one framework, operator-valued optimisation and min-max objectives are employed to address unfairness in time-series problems. This approach showcases state-of-the-art performance on the notorious COMPAS benchmark dataset, demonstrating its effectiveness in real-world scenarios. In the second framework, the challenge of lacking sensitive attributes, such as gender and race, in commonly used datasets is addressed. This issue is particularly pressing because existing algorithms in this field predominantly rely on the availability or estimations of such attributes to assess and mitigate unfairness. Here, a framework for a group-blind bias-repair is introduced, aiming to mitigate bias without relying on sensitive attributes. The efficacy of this approach is showcased through analyses conducted on the Adult Census Income dataset. Additionally, detailed algorithmic analyses for both frameworks are provided, accompanied by convergence guarantees, ensuring the robustness and reliability of the proposed methodologies.


Predicting household socioeconomic position in Mozambique using satellite and household imagery

arXiv.org Artificial Intelligence

Many studies have predicted SocioEconomic Position (SEP) for aggregated spatial units such as villages using satellite data, but SEP prediction at the household level and other sources of imagery have not been yet explored. We assembled a dataset of 975 households in a semi-rural district in southern Mozambique, consisting of self-reported asset, expenditure, and income SEP data, as well as multimodal imagery including satellite images and a ground-based photograph survey of 11 household elements. We fine-tuned a convolutional neural network to extract feature vectors from the images, which we then used in regression analyzes to model household SEP using different sets of image types. The best prediction performance was found when modeling asset-based SEP using random forest models with all image types, while the performance for expenditure- and income-based SEP was lower. Using SHAP, we observed clear differences between the images with the largest positive and negative effects, as well as identified the most relevant household elements in the predictions. Finally, we fitted an additional reduced model using only the identified relevant household elements, which had an only slightly lower performance compared to models using all images. Our results show how ground-based household photographs allow to zoom in from an area-level to an individual household prediction while minimizing the data collection effort by using explainable machine learning. The developed workflow can be potentially integrated into routine household surveys, where the collected household imagery could be used for other purposes, such as refined asset characterization and environmental exposure assessment.