Goto

Collaborating Authors

 Government


Large Language Model Lateral Spear Phishing: A Comparative Study in Large-Scale Organizational Settings

arXiv.org Artificial Intelligence

The critical threat of phishing emails has been further exacerbated by the potential of LLMs to generate highly targeted, personalized, and automated spear phishing attacks. Two critical problems concerning LLM-facilitated phishing require further investigation: 1) Existing studies on lateral phishing lack specific examination of LLM integration for large-scale attacks targeting the entire organization, and 2) Current anti-phishing infrastructure, despite its extensive development, lacks the capability to prevent LLM-generated attacks, potentially impacting both employees and IT security incident management. However, the execution of such investigative studies necessitates a real-world environment, one that functions during regular business operations and mirrors the complexity of a large organizational infrastructure. This setting must also offer the flexibility required to facilitate a diverse array of experimental conditions, particularly the incorporation of phishing emails crafted by LLMs. This study is a pioneering exploration into the use of Large Language Models (LLMs) for the creation of targeted lateral phishing emails, targeting a large tier 1 university's operation and workforce of approximately 9,000 individuals over an 11-month period. It also evaluates the capability of email filtering infrastructure to detect such LLM-generated phishing attempts, providing insights into their effectiveness and identifying potential areas for improvement. Based on our findings, we propose machine learning-based detection techniques for such emails to detect LLM-generated phishing emails that were missed by the existing infrastructure, with an F1-score of 98.96.


FactCHD: Benchmarking Fact-Conflicting Hallucination Detection

arXiv.org Artificial Intelligence

Despite their impressive generative capabilities, LLMs are hindered by fact-conflicting hallucinations in real-world applications. The accurate identification of hallucinations in texts generated by LLMs, especially in complex inferential scenarios, is a relatively unexplored area. To address this gap, we present FactCHD, a dedicated benchmark designed for the detection of fact-conflicting hallucinations from LLMs. FactCHD features a diverse dataset that spans various factuality patterns, including vanilla, multi-hop, comparison, and set operation. A distinctive element of FactCHD is its integration of fact-based evidence chains, significantly enhancing the depth of evaluating the detectors' explanations. Experiments on different LLMs expose the shortcomings of current approaches in detecting factual errors accurately. Furthermore, we introduce Truth-Triangulator that synthesizes reflective considerations by tool-enhanced ChatGPT and LoRA-tuning based on Llama2, aiming to yield more credible detection through the amalgamation of predictive results and evidence. The benchmark dataset is available at https://github.com/zjunlp/FactCHD.


Understanding the Humans Behind Online Misinformation: An Observational Study Through the Lens of the COVID-19 Pandemic

arXiv.org Artificial Intelligence

The proliferation of online misinformation has emerged as one of the biggest threats to society. Considerable efforts have focused on building misinformation detection models, still the perils of misinformation remain abound. Mitigating online misinformation and its ramifications requires a holistic approach that encompasses not only an understanding of its intricate landscape in relation to the complex issue and topic-rich information ecosystem online, but also the psychological drivers of individuals behind it. Adopting a time series analytic technique and robust causal inference-based design, we conduct a large-scale observational study analyzing over 32 million COVID-19 tweets and 16 million historical timeline tweets. We focus on understanding the behavior and psychology of users disseminating misinformation during COVID-19 and its relationship with the historical inclinations towards sharing misinformation on Non-COVID domains before the pandemic. Our analysis underscores the intricacies inherent to cross-domain misinformation, and highlights that users' historical inclination toward sharing misinformation is positively associated with their present behavior pertaining to misinformation sharing on emergent topics and beyond. This work may serve as a valuable foundation for designing user-centric inoculation strategies and ecologically-grounded agile interventions for effectively tackling online misinformation.


Uncovering local aggregated air quality index with smartphone captured images leveraging efficient deep convolutional neural network

arXiv.org Artificial Intelligence

The prevalence and mobility of smartphones make these a widely used tool for environmental health research. However, their potential for determining aggregated air quality index (AQI) based on PM2.5 concentration in specific locations remains largely unexplored in the existing literature. In this paper, we thoroughly examine the challenges associated with predicting location-specific PM2.5 concentration using images taken with smartphone cameras. The focus of our study is on Dhaka, the capital of Bangladesh, due to its significant air pollution levels and the large population exposed to it. Our research involves the development of a Deep Convolutional Neural Network (DCNN), which we train using over a thousand outdoor images taken and annotated. These photos are captured at various locations in Dhaka, and their labels are based on PM2.5 concentration data obtained from the local US consulate, calculated using the NowCast algorithm. Through supervised learning, our model establishes a correlation index during training, enhancing its ability to function as a Picture-based Predictor of PM2.5 Concentration (PPPC). This enables the algorithm to calculate an equivalent daily averaged AQI index from a smartphone image. Unlike, popular overly parameterized models, our model shows resource efficiency since it uses fewer parameters. Furthermore, test results indicate that our model outperforms popular models like ViT and INN, as well as popular CNN-based models such as VGG19, ResNet50, and MobileNetV2, in predicting location-specific PM2.5 concentration. Our dataset is the first publicly available collection that includes atmospheric images and corresponding PM2.5 measurements from Dhaka. Our codes and dataset are available at https://github.com/lepotatoguy/aqi.


Modeling and Control of a Novel Variable Stiffness Three DoFs Wrist

arXiv.org Artificial Intelligence

This study introduces an innovative design for a Variable Stiffness 3 Degrees of Freedom actuated wrist capable of actively and continuously adjusting its overall stiffness by modulating the active length of non-linear elastic elements. This modulation is akin to human muscular cocontraction and is achieved using only four motors. The mechanical configuration employed results in a compact and lightweight device with anthropomorphic characteristics, making it potentially suitable for applications such as prosthetics and humanoid robotics. This design aims to enhance performance in dynamic tasks, improve task adaptability, and ensure safety during interactions with both people and objects. The paper details the first hardware implementation of the proposed design, providing insights into the theoretical model, mechanical and electronic components, as well as the control architecture. System performance is assessed using a motion capture system. The results demonstrate that the prototype offers a broad range of motion ($[55, -45]${\deg} for flexion/extension, $\pm48${\deg} for radial/ulnar deviation, and $\pm180${\deg} for pronation/supination) while having the capability to triple its stiffness. Furthermore, following proper calibration, the wrist posture can be reconstructed through multivariate linear regression using rotational encoders and the forward kinematic model. This reconstruction achieves an average Root Mean Square Error of 6.6{\deg}, with an $R^2$ value of 0.93.


The Moderating Effect of Instant Runoff Voting

arXiv.org Artificial Intelligence

Instant runoff voting (IRV) has recently gained popularity as an alternative to plurality voting for political elections, with advocates claiming a range of advantages, including that it produces more moderate winners than plurality and could thus help address polarization. However, there is little theoretical backing for this claim, with existing evidence focused on case studies and simulations. In this work, we prove that IRV has a moderating effect relative to plurality voting in a precise sense, developed in a 1-dimensional Euclidean model of voter preferences. We develop a theory of exclusion zones, derived from properties of the voter distribution, which serve to show how moderate and extreme candidates interact during IRV vote tabulation. The theory allows us to prove that if voters are symmetrically distributed and not too concentrated at the extremes, IRV cannot elect an extreme candidate over a moderate. In contrast, we show plurality can and validate our results computationally. Our methods provide new frameworks for the analysis of voting systems, deriving exact winner distributions geometrically and establishing a connection between plurality voting and stick-breaking processes.


Ramaswamy proposes debate with Harris on AI as speculation swirls over Trump's running mate

FOX News

Vivek Ramaswamy proposed that Vice President Harris debate him on the topic of artificial intelligence, as speculation over former President Trump's choice of running mate swirls. "Kamala is in charge of AI policy right now. In a debate, I'd challenger [sic] her to see if she can spell'AI.' I'd bet on the same blank stare I got from Nikki when I asked her to name 3 provinces in eastern Ukraine," Ramaswamy wrote Tuesday night, referencing a moment from a GOP debate in December. Ramaswamy's post challenging Harris came in response to Human Events senior editor Jack Posobiec writing, "Imagine what Vivek would do to Kamala in a debate." In November, Harris notably gave a speech on the future of artifical intelligence in London, vowing that, the "United States will continue to work with the G7; the United Nations; and a diverse range of governments, from the Global North to the Global South, to promote AI safety and equity around the world."


Revealed: US police prevented from viewing many online child sexual abuse reports, lawyers say

The Guardian

Social media companies relying on artificial intelligence software to moderate their platforms are generating unviable reports on cases of child sexual abuse, preventing US police from seeing potential leads and delaying investigations of alleged predators, the Guardian can reveal. By law, US-based social media companies are required to report any child sexual abuse material detected on their platforms to the National Center for Missing & Exploited Children (NCMEC). NCMEC acts as a nationwide clearinghouse for leads about child abuse, which it forwards to the relevant law enforcement departments in the US and around the world. The organization said in its annual report that it received more than 32m reports of suspected child sexual exploitation from companies and the public in 2022, roughly 88m images, videos and other files. Meta is the largest reporter of these tips, with more than 27m, or 84%, generated by its Facebook, Instagram and WhatsApp platforms in 2022.


OpenAI Working With U.S. Military on Cybersecurity Tools

TIME - Tech

OpenAI is working with the Pentagon on a number of projects including cybersecurity capabilities, a departure from the startup's earlier ban on providing its artificial intelligence to militaries. The ChatGPT maker is developing tools with the U.S. Defense Department on open-source cybersecurity software -- collaborating with DARPA for its AI Cyber Challenge announced last year -- and has had initial talks with the US government about methods to assist with preventing veteran suicide, Anna Makanju, the company's vice president of global affairs, said in an interview at Bloomberg House at the World Economic Forum in Davos on Tuesday. The company had recently removed language in its terms of service banning its AI from "military and warfare" applications. Makanju described the decision as part of a broader update of its policies to adjust to new uses of ChatGPT and its other tools. "Because we previously had what was essentially a blanket prohibition on military, many people thought that would prohibit many of these use cases, which people think are very much aligned with what we want to see in the world," she said.


The Download: Twitter killers, and how China regulates AI

MIT Technology Review

For the better part of 17 years, the roiling, rolling, fractious, sometimes funny, sometimes horrifying, never-ever-ending global conversation had a central home: Twitter. If you wanted to know what was happening and what people were talking about right now, it was the only game in town. But then Elon Musk purchased Twitter, renamed it X, fired most of its employees, and more or less eliminated its moderation and verification systems. Many people have begun casting about for a replacement service--ideally one that is beyond any individual's control. The dream of a decentralized Twitter-like service has been around for years.