Law
Natural language processing for African languages
Recent advances in word embeddings and language models use large-scale, unlabelled data and self-supervised learning to boost NLP performance. Multilingual models, often trained on web-sourced data like Wikipedia, face challenges: few low-resource languages are included, their data is often noisy, and lack of labeled datasets makes it hard to evaluate performance outside high-resource languages like English. In this dissertation, we focus on languages spoken in Sub-Saharan Africa where all the indigenous languages in this region can be regarded as low-resourced in terms of the availability of labelled data for NLP tasks and unlabelled data found on the web. We analyse the noise in the publicly available corpora, and curate a high-quality corpus, demonstrating that the quality of semantic representations learned in word embeddings does not only depend on the amount of data but on the quality of pre-training data. We demonstrate empirically the limitations of word embeddings, and the opportunities the multilingual pre-trained language model (PLM) offers especially for languages unseen during pre-training and low-resource scenarios. We further study how to adapt and specialize multilingual PLMs to unseen African languages using a small amount of monolingual texts. To address the under-representation of the African languages in NLP research, we developed large scale human-annotated labelled datasets for 21 African languages in two impactful NLP tasks: named entity recognition and machine translation. We conduct an extensive empirical evaluation using state-of-the-art methods across supervised, weakly-supervised, and transfer learning settings.
ROSE: Toward Reality-Oriented Safety Evaluation of Large Language Models
Ding, Jiale, Zheng, Xiang, Wang, Cong, Lee, Wei-Bin, Ma, Xingjun, Jiang, Yu-Gang
As Large Language Models (LLMs) are increasingly deployed as black-box components in real-world applications, evaluating their safety-especially under adversarial prompting-has become critical. Arguably, effective safety evaluations should be adaptive, evolving with LLM capabilities, and also cover a broad spectrum of harmful topics and real-world scenarios to fully expose potential vulnerabilities. Existing manual safety benchmarks, built on handcrafted adversarial prompts, are limited by their static nature and the intensive labor required to update them, making it difficult to keep pace with rapidly advancing LLMs. In contrast, automated adversarial prompt generation offers a promising path toward adaptive evaluation. However, current methods often suffer from insufficient adversarial topic coverage (topic-level diversity) and weak alignment with real-world contexts. These shortcomings stem from the exploration-exploitation dilemma in black-box optimization and a lack of real-world contextualization, resulting in adversarial prompts that are both topically narrow and scenario-repetitive. To address these issues, we propose Reality-Oriented Safety Evaluation (ROSE), a novel framework that uses multi-objective reinforcement learning to fine-tune an adversarial LLM for generating topically diverse and contextually rich adversarial prompts. Experiments show that ROSE outperforms existing methods in uncovering safety vulnerabilities in state-of-the-art LLMs, with notable improvements in integrated evaluation metrics. We hope ROSE represents a step toward more practical and reality-oriented safety evaluation of LLMs. WARNING: This paper contains examples of potentially harmful text.
Towards Undistillable Models by Minimizing Conditional Mutual Information
Ye, Linfeng, Hamidi, Shayan Mohajer, Yang, En-hui
A deep neural network (DNN) is said to be undistillable if, when used as a black-box input-output teacher, it cannot be distilled through knowledge distillation (KD). In this case, the distilled student (referred to as the knockoff student) does not outperform a student trained independently with label smoothing (LS student) in terms of prediction accuracy. To protect intellectual property of DNNs, it is desirable to build undistillable DNNs. To this end, it is first observed that an undistillable DNN may have the trait that each cluster of its output probability distributions in response to all sample instances with the same label should be highly concentrated to the extent that each cluster corresponding to each label should ideally collapse into one probability distribution. Based on this observation and by measuring the concentration of each cluster in terms of conditional mutual information (CMI), a new training method called CMI minimized (CMIM) method is proposed, which trains a DNN by jointly minimizing the conventional cross entropy (CE) loss and the CMI values of all temperature scaled clusters across the entire temperature spectrum. The resulting CMIM model is shown, by extensive experiments, to be undistillable by all tested KD methods existing in the literature. That is, the knockoff students distilled by these KD methods from the CMIM model underperform the respective LS students. In addition, the CMIM model is also shown to performs better than the model trained with the CE loss alone in terms of their own prediction accuracy.
UN report lists companies complicit in Israel's 'genocide': Who are they?
The United Nations special rapporteur on the situation of human rights in the occupied Palestinian territory (oPt) has released a new report mapping the corporations aiding Israel in the displacement of Palestinians and its genocidal war on Gaza, in breach of international law. Francesca Albanese's latest report, which is scheduled to be presented at a news conference in Geneva on Thursday, names 48 corporate actors, including United States tech giants Microsoft, Alphabet Inc. โ Google's parent company โ and Amazon. A database of more than 1000 corporate entities was also put together as part of the investigation. "[Israel's] forever-occupation has become the ideal testing ground for arms manufacturers and Big Tech โ providing significant supply and demand, little oversight, and zero accountability โ while investors and private and public institutions profit freely," the report said. "Companies are no longer merely implicated in occupation โ they may be embedded in an economy of genocide," it said, in a reference to Israel's ongoing assault on the Gaza Strip.
AI companies start winning the copyright fight
If you need me after this newsletter publishes, I will be busy poring over photos from Jeff Bezos and Lauren Sanchez's wedding, the gaudiest and most star-studded affair to disrupt technology news this year. I found it a tacky and spectacular affair. Everyone who was anyone was there, except for Charlize Theron, who, unprompted, said on Monday: "I think we might be the only people who did not get an invite to the Bezos wedding. Judge William Alsup compared the Anthropic model's use of books to a "reader aspiring to be a writer." And the next day, Meta: The US district judge Vince Chhabria, in San Francisco, said in his decision on the Meta case that the authors had not presented enough evidence that the technology company's AI would cause "market dilution" by flooding the market with work similar to theirs. Judging by the rulings in favor of Meta and Anthropic, the authors are facing an uphill battle. Three weeks ago, Disney and NBCUniversal sued Midjourney, alleging that the ...
Senators Reject 10-Year Ban on State-Level AI Regulation, In Blow to Big Tech
Earlier in the week, Blackburn attempted to forge a compromise with Ted Cruz, who led the provision. Together, they produced a new version that reduced the ten-year ban to a five-year one, and carved out exceptions for laws related to kids' online safety and personal publicity rights. But this version of the bill was promptly excoriated by vocal coalitions in both parties. A group of 140 mostly left-leaning advocacy organizations, including Encode AI and Common Sense Media, penned an open letter arguing that this new version actually shielded tech companies from the state regulation that Blackburn was attempting to protect. "The vague standards set out in the moratorium will provide Big Tech a clear path to challenge nearly any state law in court," the letter read.
Ban on AI Regulations in Trump's Tax Bill Carries a Huge Environmental Cost
A data center for cryptocurrency mining, cloud services, and AI computing in Stutsman County, North Dakota.halbergman/Getty This story was originally published by the Guardian and is reproduced here as part of the Climate Desk collaboration. Republicans are pushing to pass a major spending bill that includes provisions to prevent states from enacting regulations on artificial intelligence. Such untamed growth in AI will take a heavy toll upon the world's dangerously overheating climate, experts have warned. About 1 billion tons of planet-heating carbon dioxide are set to be emitted in the US just from AI over the next decade if no restraints are placed on the industry's enormous electricity consumption, according to estimates by researchers at Harvard University and provided to the Guardian.
What comes next for AI copyright lawsuits?
On the other side, plaintiffs range from individual artists and authors to large companies like Getty and the New York Times. The outcomes of these cases are set to have an enormous impact on the future of AI. In effect, they will decide whether or not model makers can continue ordering up a free lunch. If not, they will need to start paying for such training data via new kinds of licensing deals--or find new ways to train their models. And that's why last week's wins for the technology companies matter. If you drill into the details, the rulings are less cut-and-dried than they seem at first.
Republicans scrap deal in 'big, beautiful bill' to lower restrictions on states' AI regulations
A deal that had been reached between Sens. Marsha Blackburn, R-Tenn., and Ted Cruz, R-Texas, over how states can regulate artificial intelligence has been pulled from President Donald Trump's "big, beautiful" bill. The collapsed agreement would have required states seeking to access hundreds of millions of dollars in AI infrastructure funding in the "big, beautiful" bill to refrain from adopting new regulations on the technology for five years, a compromise down from the original 10 years. It also included carveouts to regulate child sexual abuse material, unauthorized use of a person's likeness and other deceptive practices. Blackburn announced Monday night that she is withdrawing her support for the agreement. A deal between Sens. Marsha Blackburn and Ted Cruz over how states can regulate AI has been pulled from the "big, beautiful" bill.
Tech firms suggested placing trackers under offenders' skin at meeting with justice secretary
Tracking devices inserted under offenders' skin, robots assigned to contain prisoners and driverless vehicles used to transport them were among the measures proposed by technology companies to ministers who are gathering ideas to tackle the crisis in the UK justice system. The proposals were made at a meeting of more than two dozen tech companies in London last month, chaired by the justice secretary, Shabana Mahmood, minutes seen by the Guardian show. Amid an acute shortage of prison places and probation officers under severe strain, ministers told the companies they wanted ideas for using wearable technologies, behaviour monitoring and geolocation to create a "prison outside of prison". Those present included representatives of Google, Amazon, Microsoft and Palantir, which works closely with the US military and has contracts with the NHS. IBM and the private prison operator Serco also attended alongside tagging and biometric companies, according to a response to a freedom of information request.