Goto

Collaborating Authors

 fragmented data


Safe Training with Sensitive In-domain Data: Leveraging Data Fragmentation To Mitigate Linkage Attacks

arXiv.org Artificial Intelligence

Current text generation models are trained using real data which can potentially contain sensitive information, such as confidential patient information and the like. Under certain conditions output of the training data which they have memorised can be triggered, exposing sensitive data. To mitigate against this risk we propose a safer alternative which sees fragmented data in the form of domain-specific short phrases randomly grouped together shared instead of full texts. Thus, text fragments that could re-identify an individual cannot be reproduced by the model in one sequence, giving significant protection against linkage attacks. We fine-tune several state-of-the-art LLMs using meaningful syntactic chunks to explore their utility. In particular, we fine-tune BERT-based models to predict two cardiovascular diagnoses. Our results demonstrate the capacity of LLMs to benefit from the pre-trained knowledge and deliver classification results when fine-tuned with fragmented data comparable to fine-tuning with full training data.


How AI, data analytics could help identify insurance fraud

#artificialintelligence

People are becoming more sophisticated in perpetrating fraud--filing claims online and operating from around the world, people are approaching fraud digitally and at higher frequencies. Insurance companies have been put to the test since spring of 2020, as pandemic relief funds came into play and insurance organizations felt the changing profile of claims within personal and commercial lines. With the increasing difficulty to predict and segment claims, many organizations have found themselves falling behind, giving people a wide berth to carry out fraud without detection. Fortunately, there are new methods, using artificial intelligence, that eases the burden and helps insurance companies stay one step ahead. Historically, insurers tend to react tactically to verification and customer checks, leaving gaps in detection accuracy because of the manual nature of this process.


CenturyLink's No Sweat Approach to AI Light Reading

#artificialintelligence

"In the past, large volumes of data made us sweat". So said Pari Bajpay, vice president of Next Generation Enablement at CenturyLink, during a presentation titled "Can AI deliver its promise of a cost-effective, improved experience in telecom?" at the TM Forum's recent Digital Transformation World event in Nice. "We didn't have the networking, compute and storage capacity to cope. A lot of the data would be turned off and you would only work on the critical aspects of the data because what you had on the other end of it was humans that could not process such large volumes," noted Bajpay. However, as big data technology has matured, Bajpay and his team at CenturyLink have grappled with the issue and are now leveraging AI to extract more value from their data.


The FinTech Outlook for 2018

#artificialintelligence

Getting down to business with Artificial Intelligence (AI): several of the large banks like JPMorgan and UBS were doing interesting things with AI in 2017, but I think this will be a lot more pervasive and recognised across all banks in 2018. This is primarily for compliance and risk, as AI can develop and apply complex rules across all business processes in real-time, and when most banks have 1 in 3 staff checking compliance, it makes absolute sense to replace them with learning software. Rationalising and cleansing core data structures: many banks have built their core operations on fragmented systems aligned to products. This has distributed customer data across multiple platforms, and banks recognise that they cannot use AI effectively on data spread across the business. As a result, many will develop strategies for building an Enterprise Data Architecture in 2018 which rationalises and cleanses their fragmented data stores.