Government
Professional Certification Benchmark Dataset: The First 500 Jobs For Large Language Models
The research creates a professional certification survey to test large language models and evaluate their employable skills. It compares the performance of two AI models, GPT-3 and Turbo-GPT3.5, on a benchmark dataset of 1149 professional certifications, emphasizing vocational readiness rather than academic performance. GPT-3 achieved a passing score (>70% correct) in 39% of the professional certifications without fine-tuning or exam preparation. The models demonstrated qualifications in various computer-related fields, such as cloud and virtualization, business analytics, cybersecurity, network setup and repair, and data analytics. Turbo-GPT3.5 scored 100% on the valuable Offensive Security Certified Professional (OSCP) exam. The models also displayed competence in other professional domains, including nursing, licensed counseling, pharmacy, and teaching. Turbo-GPT3.5 passed the Financial Industry Regulatory Authority (FINRA) Series 6 exam with a 70% grade without preparation. Interestingly, Turbo-GPT3.5 performed well on customer service tasks, suggesting potential applications in human augmentation for chatbots in call centers and routine advice services. The models also score well on sensory and experience-based tests such as wine sommelier, beer taster, emotional quotient, and body language reader. The OpenAI model improvement from Babbage to Turbo resulted in a median 60% better-graded performance in less than a few years. This progress suggests that focusing on the latest model's shortcomings could lead to a highly performant AI capable of mastering the most demanding professional certifications. We open-source the benchmark to expand the range of testable professional skills as the models improve or gain emergent capabilities.
Wasserstein multivariate auto-regressive models for modeling distributional time series and its application in graph learning
We propose a new auto-regressive model for the statistical analysis of multivariate distributional time series. The data of interest consist of a collection of multiple series of probability measures supported over a bounded interval of the real line, and that are indexed by distinct time instants. The probability measures are modelled as random objects in the Wasserstein space. We establish the auto-regressive model in the tangent space at the Lebesgue measure by first centering all the raw measures so that their Fr\'echet means turn to be the Lebesgue measure. Using the theory of iterated random function systems, results on the existence, uniqueness and stationarity of the solution of such a model are provided. We also propose a consistent estimator for the model coefficient. In addition to the analysis of simulated data, the proposed model is illustrated with two real data sets made of observations from age distribution in different countries and bike sharing network in Paris. Finally, due to the positive and boundedness constraints that we impose on the model coefficients, the proposed estimator that is learned under these constraints, naturally has a sparse structure. The sparsity allows furthermore the application of the proposed model in learning a graph of temporal dependency from the multivariate distributional time series.
Leveraging Semantic Relationships to Prioritise Indicators of Compromise in Additive Manufacturing Systems
Kumar, Mahender, Epiphaniou, Gregory, Maple, Carsten
Additive manufacturing (AM) offers numerous benefits, such as manufacturing complex and customised designs quickly and cost-effectively, reducing material waste, and enabling on-demand production. However, several security challenges are associated with AM, making it increasingly attractive to attackers ranging from individual hackers to organised criminal gangs and nation-state actors. This paper addresses the cyber risk in AM to attackers by proposing a novel semantic-based threat prioritisation system for identifying, extracting and ranking indicators of compromise (IOC). The system leverages the heterogeneous information networks (HINs) that automatically extract high-level IOCs from multi-source threat text and identifies semantic relations among the IOCs. It models IOCs with a HIN comprising different meta-paths and meta-graphs to depict semantic relations among diverse IOCs. We introduce a domain-specific recogniser that identifies IOCs in three domains: organisation-specific, regional source-specific, and regional target-specific. A threat assessment uses similarity measures based on meta-paths and meta-graphs to assess semantic relations among IOCs. It prioritises IOCs by measuring their severity based on the frequency of attacks, IOC lifetime, and exploited vulnerabilities in each domain.
Symbolic Regression on FPGAs for Fast Machine Learning Inference
Tsoi, Ho Fung, Pol, Adrian Alan, Loncar, Vladimir, Govorkova, Ekaterina, Cranmer, Miles, Dasu, Sridhara, Elmer, Peter, Harris, Philip, Ojalvo, Isobel, Pierini, Maurizio
The high-energy physics community is investigating the feasibility of deploying machine-learning-based solutions on Field-Programmable Gate Arrays (FPGAs) to improve physics sensitivity while meeting data processing latency limitations. In this contribution, we introduce a novel end-to-end procedure that utilizes a machine learning technique called symbolic regression (SR). It searches equation space to discover algebraic relations approximating a dataset. We use PySR (software for uncovering these expressions based on evolutionary algorithm) and extend the functionality of hls4ml (a package for machine learning inference in FPGAs) to support PySR -generated expressions for resource-constrained production environments. Deep learning models often optimise the top metric by pinning the network size because vast hyperparameter space prevents extensive neural architecture search. Conversely, SR selects a set of models on the Pareto front, which allows for optimising the performanceresource tradeoff directly. By embedding symbolic forms, our implementation can dramatically reduce the computational resources needed to perform critical tasks. We validate our procedure on a physics benchmark: multiclass classification of jets produced in simulated proton-proton collisions at the CERN Large Hadron Collider, and show that we approximate a 3-layer neural network with an inference model that has as low as 5 ns execution time (a reduction by a factor of 13) and over 90% approximation accuracy.
Toucha11y: Making Inaccessible Public Touchscreens Accessible
Li, Jiasheng, Yan, Zeyu, Shah, Arush, Lazar, Jonathan, Peng, Huaishu
Despite their growing popularity, many public kiosks with touchscreens are inaccessible to blind people. Toucha11y is a working prototype that allows blind users to use existing inaccessible touchscreen kiosks independently and with little effort. Toucha11y consists of a mechanical bot that can be instrumented to an arbitrary touchscreen kiosk by a blind user and a companion app on their smartphone. The bot, once attached to a touchscreen, will recognize its content, retrieve the corresponding information from a database, and render it on the user's smartphone. As a result, a blind person can use the smartphone's built-in accessibility features to access content and make selections. The mechanical bot will detect and activate the corresponding touchscreen interface. We present the system design of Toucha11y along with a series of technical evaluations. Through a user study, we found out that Toucha11y could help blind users operate inaccessible touchscreen devices.
Generalization Guarantees for Multi-item Profit Maximization: Pricing, Auctions, and Randomized Mechanisms
Balcan, Maria-Florina, Sandholm, Tuomas, Vitercik, Ellen
We study multi-item profit maximization when there is an underlying distribution over buyers' values. In practice, a full description of the distribution is typically unavailable, so we study the setting where the mechanism designer only has samples from the distribution. If the designer uses the samples to optimize over a complex mechanism class -- such as the set of all multi-item, multi-buyer mechanisms -- a mechanism may have high average profit over the samples but low expected profit. This raises the central question of this paper: how many samples are sufficient to ensure that a mechanism's average profit is close to its expected profit? To answer this question, we uncover structure shared by many pricing, auction, and lottery mechanisms: for any set of buyers' values, profit is piecewise linear in the mechanism's parameters. Using this structure, we prove new bounds for mechanism classes not yet studied in the sample-based mechanism design literature and match or improve over the best-known guarantees for many classes.
If a Press Release Says "Artificial Intelligence," There's a Good Chance It's Meaningless
This article is from Big Technology, a newsletter by Alex Kantrowitz. You can already see the machine at work. Corporations, politicians, threadbois, and "thought leaders" are probing and prodding, searching desperately for ways to use surging curiosity about all things artificial intelligence to mask problems, gain favor with the public, and monetize attention. This A.I.-PR industrial complex is growing larger and worse than its predecessors--even crypto!--because the technology is making anything seem possible. With so much opportunity, vacuousness fills the gaps, and exploitation follows.
'We've discovered the secret of immortality. The bad news is it's not for us': why the godfather of AI fears for humanity
The first thing Geoffrey Hinton says when we start talking, and the last thing he repeats before I turn off my recorder, is that he left Google, his employer of the past decade, on good terms. "I have no objection to what Google has done or is doing, but obviously the media would love to spin me as'a disgruntled Google employee'. It's an important clarification to make, because it's easy to conclude the opposite. After all, when most people calmly describe their former employer as being one of a small group of companies charting a course that is alarmingly likely to wipe out humanity itself, they do so with a sense of opprobrium. But to listen to Hinton, we're about to sleepwalk towards an existential threat to civilisation without anyone involved acting maliciously at all. Known as one of three "godfathers of AI", in 2018 Hinton won the ACM Turing award โ the Nobel prize of computer scientists for his work on "deep learning". A cognitive psychologist and computer scientist by training, he wasn't motivated by a desire to radically improve technology: instead, it was to understand more about ourselves. "For the last 50 years, I've been trying to make computer models that can learn stuff a bit like the way the brain learns it, in order to understand better how the brain is learning things," he tells me when we meet in his sister's house in north London, where he is staying (he usually resides in Canada). Looming slightly over me โ he prefers to talk standing up, he says โ the tone is uncannily reminiscent of a university tutorial, as the 75-year-old former professor explains his research history, and how it has inescapably led him to the conclusion that we may be doomed. In trying to model how the human brain works, Hinton found himself one of the leaders in the field of "neural networking", an approach to building computer systems that can learn from data and experience. Until recently, neural nets were a curiosity, requiring vast computer power to perform simple tasks worse than other approaches. But in the last decade, as the availability of processing power and vast datasets has exploded, the approach Hinton pioneered has ended up at the centre of a technological revolution. "In trying to think about how the brain could implement the algorithm behind all these models, I decided that maybe it can't โ and maybe these big models are actually much better than the brain," he says. A "biological intelligence" such as ours, he says, has advantages. It runs at low power, "just 30 watts, even when you're thinking", and "every brain is a bit different". That means we learn by mimicking others. But that approach is "very inefficient" in terms of information transfer. Digital intelligences, by contrast, have an enormous advantage: it's trivial to share information between multiple copies. "You pay an enormous cost in terms of energy, but when one of them learns something, all of them know it, and you can easily store more copies.
Why was the Kremlin attacked? In Russia, it depends who you ask
An apparent drone attack on the Kremlin this week has sparked fears of an escalation in Russia's brutal war in Ukraine. On Wednesday night, two remotely-operated devices flew towards the domed roof of the Kremlin before being shot down by Russian air defences, exploding but harming no one. After the incident, Moscow Mayor Sergey Sobyanin declared that flying drones by private citizens was now banned in Moscow. Russia said the United States masterminded the attack, claiming Ukraine carried it out. Washington and Kyiv have denied responsibility, insisting that Ukraine's war efforts are purely defensive.
The Thorny Art of Deepfake Labeling
Last week, the Republican National Committee put out a video advertisement against Biden, which featured a small disclaimer in the top left of the frame: "Built entirely with AI imagery." Critics questioned the diminished size of the disclaimer and suggested its limited value, particularly because the ad marks the first substantive use of AI in political attack advertising. As AI-generated media become more mainstream, many have argued that text-based labels, captions, and watermarks are crucial for transparency. But do these labels actually work? For a label to work, it needs to be legible.