Government
Fine-Tuning LLMs with Noisy Data for Political Argument Generation and Post Guidance
Churina, Svetlana, Jaidka, Kokil
The incivility in social media discourse complicates the deployment of automated text generation models for politically sensitive content. Fine-tuning and prompting strategies are critical, but underexplored, solutions to mitigate toxicity in such contexts. This study investigates the fine-tuning and prompting effects on GPT-3.5 Turbo using subsets of the CLAPTON dataset of political discussion posts, comprising Twitter and Reddit data labeled for their justification, reciprocity and incivility. Fine-tuned models on Reddit data scored highest on discussion quality, while combined noisy data led to persistent toxicity. Prompting strategies reduced specific toxic traits, such as personal attacks, but had limited broader impact. The findings emphasize that high-quality data and well-crafted prompts are essential to reduce incivility and improve rhetorical quality in automated political discourse generation.
SMI-Editor: Edit-based SMILES Language Model with Fragment-level Supervision
Zheng, Kangjie, Liang, Siyue, Yang, Junwei, Feng, Bin, Liu, Zequn, Ju, Wei, Xiao, Zhiping, Zhang, Ming
SMILES, a crucial textual representation of molecular structures, has garnered significant attention as a foundation for pre-trained language models (LMs). However, most existing pre-trained SMILES LMs focus solely on the single-token level supervision during pre-training, failing to fully leverage the substructural information of molecules. This limitation makes the pre-training task overly simplistic, preventing the models from capturing richer molecular semantic information. Moreover, during pre-training, these SMILES LMs only process corrupted SMILES inputs, never encountering any valid SMILES, which leads to a train-inference mismatch. To address these challenges, we propose SMI-Editor, a novel edit-based pre-trained SMILES LM. SMI-Editor disrupts substructures within a molecule at random and feeds the resulting SMILES back into the model, which then attempts to restore the original SMILES through an editing process. This approach not only introduces fragment-level training signals, but also enables the use of valid SMILES as inputs, allowing the model to learn how to reconstruct complete molecules from these incomplete structures. As a result, the model demonstrates improved scalability and an enhanced ability to capture fragment-level molecular information. Experimental results show that SMI-Editor achieves state-of-the-art performance across multiple downstream molecular tasks, and even outperforming several 3D molecular representation models.
REGE: A Method for Incorporating Uncertainty in Graph Embeddings
Shafi, Zohair, Savcisens, Germans, Eliassi-Rad, Tina
Machine learning models for graphs in real-world applications are prone to two primary types of uncertainty: (1) those that arise from incomplete and noisy data and (2) those that arise from uncertainty of the model in its output. These sources of uncertainty are not mutually exclusive. Additionally, models are susceptible to targeted adversarial attacks, which exacerbate both of these uncertainties. In this work, we introduce Radius Enhanced Graph Embeddings (REGE), an approach that measures and incorporates uncertainty in data to produce graph embeddings with radius values that represent the uncertainty of the model's output. REGE employs curriculum learning to incorporate data uncertainty and conformal learning to address the uncertainty in the model's output. In our experiments, we show that REGE's graph embeddings perform better under adversarial attacks by an average of 1.5% (accuracy) against state-of-the-art methods.
Strategizing Equitable Transit Evacuations: A Data-Driven Reinforcement Learning Approach
Tang, Fang, Wang, Han, Monache, Maria Laura Delle
As natural disasters become increasingly frequent, the need for efficient and equitable evacuation planning has become more critical. This paper proposes a data-driven, reinforcement learning-based framework to optimize bus-based evacuations with an emphasis on improving both efficiency and equity. We model the evacuation problem as a Markov Decision Process solved by reinforcement learning, using real-time transit data from General Transit Feed Specification and transportation networks extracted from OpenStreetMap. The reinforcement learning agent dynamically reroutes buses from their scheduled location to minimize total passengers' evacuation time while prioritizing equity-priority communities. Simulations on the San Francisco Bay Area transportation network indicate that the proposed framework achieves significant improvements in both evacuation efficiency and equitable service distribution compared to traditional rule-based and random strategies. These results highlight the potential of reinforcement learning to enhance system performance and urban resilience during emergency evacuations, offering a scalable solution for real-world applications in intelligent transportation systems.
KITE-DDI: A Knowledge graph Integrated Transformer Model for accurately predicting Drug-Drug Interaction Events from Drug SMILES and Biomedical Knowledge Graph
Tamir, Azwad, Yuan, Jiann-Shiun
It is a common practice in modern medicine to prescribe multiple medications simultaneously to treat diseases. However, these medications could have adverse reactions between them, known as Drug-Drug Interactions (DDI), which have the potential to cause significant bodily injury and could even be fatal. Hence, it is essential to identify all the DDI events before prescribing multiple drugs to a patient. Most contemporary research for predicting DDI events relies on either information from Biomedical Knowledge graphs (KG) or drug SMILES, with very few managing to merge data from both to make predictions. While others use heuristic algorithms to extract features from SMILES and KGs, which are then fed into a Deep Learning framework to generate output. In this study, we propose a KG-integrated Transformer architecture to generate an end-to-end fully automated Machine Learning pipeline for predicting DDI events with high accuracy. The algorithm takes full-scale molecular SMILES sequences of a pair of drugs and a biomedical KG as input and predicts the interaction between the two drugs with high precision. The results show superior performance in two different benchmark datasets compared to existing state-of-the-art models especially when the test and training sets contain distinct sets of drug molecules. This demonstrates the strong generalization of the proposed model, indicating its potential for DDI event prediction for newly developed drugs. The model does not depend on heuristic models for generating embeddings and has a minimal number of hyperparameters, making it easy to use while demonstrating outstanding performance in low-data scenarios.
Designing Domain-Specific Large Language Models: The Critical Role of Fine-Tuning in Public Opinion Simulation
Large language models (LLMs) have transformed natural language processing, yet face challenges in specialized tasks such as simulating opinions on environmental policies. This paper introduces a novel fine-tuning approach that integrates socio-demographic data from the UK Household Longitudinal Study, uniquely using profiling factors, such as age, gender, income, education, and region. This method enhances the accuracy and representation of generated views. By emulating diverse synthetic profiles, the fine-tuned models significantly outperform pre-trained counterparts, achieving measurable improvements in capturing demographic nuances. Evaluation metrics, including Chi-Squared, Cosine Similarity, Jaccard Index, and KL-divergence, reveal a strong alignment between synthetic and real-world opinions. This work demonstrates the potential of fine-tuned LLMs tailored to societal contexts to enable more ethical and precise policy simulations. Its broader implications include deploying LLMs in domains like healthcare and education, fostering inclusive and data-driven decision-making in both research and practice.
If you're really bored, X's Grok AI chatbot is now free to use
Is your weekend a bit bare-bones? Here's something that could entertain you for a minute or two. The chatbot Grok-2 is now free for everyone to fool around with on X. We knew this was coming and, well, now it's here. There are some limitations for those who don't want to plunk down 8 (or more) each month for X Premium. The free tier only allows for ten messages in each two-hour period.
Mystery of bizarre drones over New Jersey deepens after new footage of UFOs emerge
New footage of multiple eerie'triangle' craft flying above New Jersey has only compounded the mystery for locals. At least five or possibly six of the unidentified drones were captured in the new, 50-second cell phone video, which one commenter declared was'the clearest video yet.' One drone, heard roaring in the skies as it moved through the darkness, appeared to have a cluster of white lights on its underbelly and red lights blinking at the tips of its wings and tail. Another drone came into frame that resembled a classic'black triangle' UFO or the triangular TR-3B, which beamed bright white lights from its nose, wingtips and tail. Since mid-November, a wave of unexplained drone sightings above central Jersey has left both law enforcement and the general public watching the skies, hunting for clues on what these mysterious night flights might be.
Revealed: bias found in AI system used to detect UK benefits fraud
An artificial intelligence system used by the UK government to detect welfare fraud is showing bias according to people's age, disability, marital status and nationality, the Guardian can reveal. An internal assessment of a machine-learning programme used to vet thousands of claims for universal credit payments across England found it incorrectly selected people from some groups more than others when recommending whom to investigate for possible fraud. The admission was made in documents released under the Freedom of Information Act by the Department for Work and Pensions (DWP). The "statistically significant outcome disparity" emerged in a "fairness analysis" of the automated system for universal credit advances carried out in February this year. The emergence of the bias comes after the DWP this summer claimed the AI system "does not present any immediate concerns of discrimination, unfair treatment or detrimental impact on customers". This assurance came in part because the final decision on whether a person gets a welfare payment is still made by a human, and officials believe the continued use of the system – which is attempting to help cut an estimated 8bn a year lost in fraud and error – is "reasonable and proportionate".
New Jersey leaders speak to DHS as unusual drone sightings now also reported over New York
Officials are still investigating unusual drone activity that has been reported in recent weeks in New Jersey. The FAA set temporary restrictions above Trump National Golf Club in Bedminster in response. New Jersey Gov. Phil Murphy said he spoke with state and federal officials about unusual drone activity in parts of the region, including the vicinity of President-elect Trump's Bedminster golf club, but stressed there was no threat to public safety. In a Thursday post on X, Murphy said he convened a briefing with Homeland Security Secretary Alejandro Mayorkas, senior officials from DHS, the state police and state Homeland Security and Preparedness, as well as New Jersey's congressional delegation. "We are actively monitoring the situation and in close coordination with our federal and law enforcement partners on this matter," he wrote.