AITopics | metadata type

Collaborating Authors

metadata type

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

If you are looking for an answer to the question What is Artificial Intelligence? and you only have a minute, then here's the definition the Association for the Advancement of Artificial Intelligence offers on its home page: "the scientific understanding of the mechanisms underlying thought and intelligent behavior and their embodiment in machines."

However, if you are fortunate enough to have more than a minute, then please get ready to embark upon an exciting journey exploring AI (but beware, it could last a lifetime) …

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining

Fan, Dongyang, Hashemi, Diba, Karimireddy, Sai Praneeth, Jaggi, Martin

arXiv.org Artificial IntelligenceNov-27-2025

Incorporating metadata in Large Language Models (LLMs) pretraining has recently emerged as a promising approach to accelerate training. However prior work highlighted only one useful signal-URLs, leaving open the question of whether other forms of metadata could yield greater benefits. In this study, we investigate a wider range of metadata types and find other types of metadata, such as fine-grained indicators of document quality that can also accelerate pretraining when prepended. We identify a common feature among effective metadata: they encode information at a finer granularity. We further introduce metadata appending as a means of improving training efficiency, where predicting an appropriate metadata as auxiliary task can help speed up pretraining. In addition, learnable meta-tokens trained with masked loss can recover part of the speedup by inducing quality-aware latent structure. Using probing, we analyze latent representations to understand how metadata shapes learning. Together, these results yield practical guidelines for integrating metadata to improve both the efficiency and effectiveness of LLM pretraining.

artificial intelligence, large language model, natural language, (18 more...)

arXiv.org Artificial Intelligence

2511.21613

Country: North America > United States (0.28)

Genre: Research Report > New Finding (0.48)

Technology: Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)

Add feedback

A Preliminary Study of RAG for Taiwanese Historical Archives

Lin, Claire, Feng, Bo-Han, Chen, Xuanjun, Yang, Te-Lun, Lee, Hung-yi, Jang, Jyh-Shing Roger

arXiv.org Artificial IntelligenceNov-12-2025

Retrieval-Augmented Generation (RAG) has emerged as a promising approach for knowledge-intensive tasks. However, few studies have examined RAG for Taiwanese Historical Archives. In this paper, we present an initial study of a RAG pipeline applied to two historical Traditional Chinese datasets, Fort Zeelandia and the Taiwan Provincial Council Gazette, along with their corresponding open-ended query sets. We systematically investigate the effects of query characteristics and metadata integration strategies on retrieval quality, answer generation, and the performance of the overall system. The results show that early-stage metadata integration enhances both retrieval and answer accuracy while also revealing persistent challenges for RAG systems, including hallucinations during generation and difficulties in handling temporal or multi-hop historical queries.

large language model, machine learning, natural language, (20 more...)

arXiv.org Artificial Intelligence

2511.07445

Country:

North America (0.46)
Asia > Taiwan (0.26)

Genre:

Research Report > New Finding (0.88)
Research Report > Experimental Study (0.68)

Technology:

Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.91)

Add feedback

AI-enabled Automation for Completeness Checking of Privacy Policies

Amaral, Orlando, Abualhaija, Sallam, Torre, Damiano, Sabetzadeh, Mehrdad, Briand, Lionel C.

arXiv.org Artificial IntelligenceJun-10-2021

Technological advances in information sharing have raised concerns about data protection. Privacy policies contain privacy-related requirements about how the personal data of individuals will be handled by an organization or a software system (e.g., a web service or an app). In Europe, privacy policies are subject to compliance with the General Data Protection Regulation (GDPR). A prerequisite for GDPR compliance checking is to verify whether the content of a privacy policy is complete according to the provisions of GDPR. Incomplete privacy policies might result in large fines on violating organization as well as incomplete privacy-related software specifications. Manual completeness checking is both time-consuming and error-prone. In this paper, we propose AI-based automation for the completeness checking of privacy policies. Through systematic qualitative methods, we first build two artifacts to characterize the privacy-related provisions of GDPR, namely a conceptual model and a set of completeness criteria. Then, we develop an automated solution on top of these artifacts by leveraging a combination of natural language processing and supervised machine learning. Specifically, we identify the GDPR-relevant information content in privacy policies and subsequently check them against the completeness criteria. To evaluate our approach, we collected 234 real privacy policies from the fund industry. Over a set of 48 unseen privacy policies, our approach detected 300 of the total of 334 violations of some completeness criteria correctly, while producing 23 false positives. The approach thus has a precision of 92.9% and recall of 89.8%. Compared to a baseline that applies keyword search only, our approach results in an improvement of 24.5% in precision and 38% in recall.

metadata type, personal data, privacy policy, (15 more...)

arXiv.org Artificial Intelligence

2106.05688

Country:

Europe > Jersey (0.14)
South America > Argentina (0.04)
Oceania > New Zealand (0.04)
(15 more...)

Genre: Research Report (1.00)

Industry:

Law (1.00)
Information Technology > Security & Privacy (1.00)

Technology:

Information Technology > Security & Privacy (1.00)
Information Technology > Artificial Intelligence > Natural Language > Text Processing (1.00)
Information Technology > Artificial Intelligence > Machine Learning > Statistical Learning (0.93)
(2 more...)

Add feedback

A Map Equation with Metadata: Varying the Role of Attributes in Community Detection

Emmons, Scott, Mucha, Peter J.

arXiv.org Machine LearningOct-24-2018

As the No Free Lunch theorem formally states [1], algorithms for detecting communities in networks must make tradeoffs. In this work, we present a method for using metadata to inform tradeoff decisions. We extend the content map equation, which adds metadata entropy to the traditional map equation, by introducing a tuning parameter allowing for explicit specification of the metadata's relative importance in assigning community labels. On synthetic networks, we show how tuning for node metadata relates to the detectability limit, and on empirical networks, we show how increased tuning for node metadata yields increased mutual information with the metadata at a cost in the traditional map equation. Our tuning parameter, like the focusing knob of a microscope, allows users to "zoom in" and "zoom out" on communities with varying levels of focus on the metadata.

artificial intelligence, data mining, metadata, (19 more...)

arXiv.org Machine Learning

1810.10433

Country: North America > United States > North Carolina (0.28)

Genre: Research Report (1.00)

Industry:

Law (0.94)
Education > Educational Setting > Higher Education (0.46)
Education > Educational Setting > K-12 Education (0.30)

Technology:

Information Technology > Artificial Intelligence (0.68)
Information Technology > Data Science > Data Mining (0.52)

Add feedback