Goto

Collaborating Authors

 esco


Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies

arXiv.org Artificial Intelligence

This study investigates the potential of language models to improve the classification of labor market information by linking job vacancy texts to two major European frameworks: the European Skills, Competences, Qualifications and Occupations (ESCO) taxonomy and the European Qualifications Framework (EQF). We examine and compare two prominent methodologies from the literature: Sentence Linking and Entity Linking. In support of ongoing research, we release an open-source tool, incorporating these two methodologies, designed to facilitate further work on labor classification and employment discourse. To move beyond surface-level skill extraction, we introduce two annotated datasets specifically aimed at evaluating how occupations and qualifications are represented within job vacancy texts. Additionally, we examine different ways to utilize generative large language models for this task. Our findings contribute to advancing the state of the art in job entity extraction and offer computational infrastructure for examining work, skills, and labor market narratives in a digitally mediated economy. Our code is made publicly available: https://github.com/tabiya-tech/tabiya-livelihoods-classifier


Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Ranker

arXiv.org Artificial Intelligence

Workforce transformation across diverse industries has driven an increased demand for specialized natural language processing capabilities. Nevertheless, tasks derived from work-related contexts inherently reflect real-world complexities, characterized by long-tailed distributions, extreme multi-label target spaces, and scarce data availability. The rise of generalist embedding models prompts the question of their performance in the work domain, especially as progress in the field has focused mainly on individual tasks. To this end, we introduce WorkBench, the first unified evaluation suite spanning six work-related tasks formulated explicitly as ranking problems, establishing a common ground for multi-task progress. Based on this benchmark, we find significant positive cross-task transfer, and use this insight to compose task-specific bipartite graphs from real-world data, synthetically enriched through grounding. This leads to Unified Work Em-beddings (UWE), a task-agnostic bi-encoder that exploits our training-data structure with a many-to-many InfoNCE objective, and leverages token-level embeddings with task-agnostic soft late interaction. UWE demonstrates zero-shot ranking performance on unseen target spaces in the work domain, enables low-latency inference by caching the task target space embeddings, and shows significant gains in macro-averaged MAP and RP@10 over generalist embedding models.


Lifelong Event Detection with Embedding Space Separation and Compaction

arXiv.org Artificial Intelligence

To mitigate forgetting, existing lifelong event detection methods typically maintain a memory module and replay the stored memory data during the learning of a new task. However, the simple combination of memory data and new-task samples can still result in substantial forgetting of previously acquired knowledge, which may occur due to the potential overlap between the feature distribution of new data and the previously learned embedding space. Moreover, the model suffers from overfitting on the few memory samples rather than effectively remembering learned patterns. To address the challenges of forgetting and overfitting, we propose a novel method based on embedding space separation and compaction. Our method alleviates forgetting of previously learned tasks by forcing the feature distribution of new data away from the previous embedding space. It also mitigates overfitting by a memory calibration mechanism that encourages memory data to be close to its prototype to enhance intra-class compactness. In addition, the learnable parameters of the new task are initialized by drawing upon acquired knowledge from the previously learned task to facilitate forward knowledge transfer. With extensive experiments, we demonstrate that our method can significantly outperform previous state-of-the-art approaches.


JOBSKAPE: A Framework for Generating Synthetic Job Postings to Enhance Skill Matching

arXiv.org Artificial Intelligence

Recent approaches in skill matching, employing synthetic training data for classification or similarity model training, have shown promising results, reducing the need for time-consuming and expensive annotations. However, previous synthetic datasets have limitations, such as featuring only one skill per sentence and generally comprising short sentences. In this paper, we introduce JobSkape, a framework to generate synthetic data that tackles these limitations, specifically designed to enhance skill-to-taxonomy matching. Within this framework, we create SkillSkape, a comprehensive open-source synthetic dataset of job postings tailored for skill-matching tasks. We introduce several offline metrics that show that our dataset resembles real-world data. Additionally, we present a multi-step pipeline for skill extraction and matching tasks using large language models (LLMs), benchmarking against known supervised methodologies. We outline that the downstream evaluation results on real-world data can beat baselines, underscoring its efficacy and adaptability.


Entity Linking in the Job Market Domain

arXiv.org Artificial Intelligence

In Natural Language Processing, entity linking (EL) has centered around Wikipedia, but yet remains underexplored for the job market domain. Disambiguating skill mentions can help us get insight into the current labor market demands. In this work, we are the first to explore EL in this domain, specifically targeting the linkage of occupational skills to the ESCO taxonomy (le Vrang et al., 2014). Previous efforts linked coarse-grained (full) sentences to a corresponding ESCO skill. In this work, we link more fine-grained span-level mentions of skills. We tune two high-performing neural EL models, a bi-encoder (Wu et al., 2020) and an autoregressive model (Cao et al., 2021), on a synthetically generated mention--skill pair dataset and evaluate them on a human-annotated skill-linking benchmark. Our findings reveal that both models are capable of linking implicit mentions of skills to their correct taxonomy counterparts. Empirically, BLINK outperforms GENRE in strict evaluation, but GENRE performs better in loose evaluation (accuracy@$k$).


Beyond experts: jobs, tasks, and skills for a data driven Future of Work ZDNet

#artificialintelligence

The World Economic Forum (WEF) is upon us this week, and the Future of Work is one of its key themes. This is a good opportunity to catch up on the trends unfolding in this domain right now, and to ponder on the insights of the people taking note and shaping this discussion. Automation and AI is part of this discussion as well, with the jury still out as to how exactly this will shape labor, workforce dynamics, and workplace transformation among others. Based on the WEF's latest report on the Future of Jobs, we highlight the major forces at play today. We discuss how these effect the technology behind the job market with Panos Alexopoulos, Head of Ontology at Textkernel, a Careerbuilder company.


New Company Uses Artificial Intelligence To Sell You Cheaper Power

#artificialintelligence

A new startup launched today that seeks to streamline electricity purchases with the use of a peer-to-peer network that relies on artificial intelligence and other technology like high-frequency trading to balance supply and demand in real time. The idea behind the company, named Drift, is to remove inefficiencies from the current wholesale power market structure in deregulated service areas and pass the savings on to consumers. Drift is initially rolling out its service in New York City and plans to expand to other parts of the country this year. The company will operate like a mini Independent System Operator (ISO) for a distributed energy platform. Drift has nearly 3,000 independent power producers in its network that range from hydroelectric dams, to solar-plus-storage projects, wind farms and large commercial buildings.