Large Language Model
Multi-Modal CLIP-Informed Protein Editing
Yin, Mingze, Zhou, Hanjing, Zhu, Yiheng, Lin, Miao, Wu, Yixuan, Wu, Jialu, Xu, Hongxia, Hsieh, Chang-Yu, Hou, Tingjun, Chen, Jintai, Wu, Jian
Proteins govern most biological functions essential for life, but achieving controllable protein discovery and optimization remains challenging. Recently, machine learning-assisted protein editing (MLPE) has shown promise in accelerating optimization cycles and reducing experimental workloads. However, current methods struggle with the vast combinatorial space of potential protein edits and cannot explicitly conduct protein editing using biotext instructions, limiting their interactivity with human feedback. To fill these gaps, we propose a novel method called ProtET for efficient CLIP-informed protein editing through multi-modality learning. Our approach comprises two stages: in the pretraining stage, contrastive learning aligns protein-biotext representations encoded by two large language models (LLMs), respectively. Subsequently, during the protein editing stage, the fused features from editing instruction texts and original protein sequences serve as the final editing condition for generating target protein sequences. Comprehensive experiments demonstrated the superiority of ProtET in editing proteins to enhance human-expected functionality across multiple attribute domains, including enzyme catalytic activity, protein stability and antibody specific binding ability. And ProtET improves the state-of-the-art results by a large margin, leading to significant stability improvements of 16.67% and 16.90%. This capability positions ProtET to advance real-world artificial protein editing, potentially addressing unmet academic, industrial, and clinical needs.
Why Misinformation is Created? Detecting them by Integrating Intent Features
Wang, Bing, Li, Ximing, Li, Changchun, Fu, Bo, Pei, Songwen, Wang, Shengsheng
Various social media platforms, e.g., Twitter and Reddit, allow people to disseminate a plethora of information more efficiently and conveniently. However, they are inevitably full of misinformation, causing damage to diverse aspects of our daily lives. To reduce the negative impact, timely identification of misinformation, namely Misinformation Detection (MD), has become an active research topic receiving widespread attention. As a complex phenomenon, the veracity of an article is influenced by various aspects. In this paper, we are inspired by the opposition of intents between misinformation and real information. Accordingly, we propose to reason the intent of articles and form the corresponding intent features to promote the veracity discrimination of article features. To achieve this, we build a hierarchy of a set of intents for both misinformation and real information by referring to the existing psychological theories, and we apply it to reason the intent of articles by progressively generating binary answers with an encoder-decoder structure. We form the corresponding intent features and integrate it with the token features to achieve more discriminative article features for MD. Upon these ideas, we suggest a novel MD method, namely Detecting Misinformation by Integrating Intent featuRes (DM-INTER). To evaluate the performance of DM-INTER, we conduct extensive experiments on benchmark MD datasets. The experimental results validate that DM-INTER can outperform the existing baseline MD methods.
Towards the Terminator Economy: Assessing Job Exposure to AI through LLMs
Colombo, Emilio, Mercorio, Fabio, Mezzanzanica, Mario, Serino, Antonio
The spread and rapid development of AI-related technologies are influencing many aspects of our daily lives, from social to educational, including the labour market. Many researchers have been highlighting the key role AI and technologies play in reshaping jobs and their related tasks, either by automating or enhancing human capabilities in the workplace. Can we estimate if, and to what extent, jobs and related tasks are exposed to the risk of being automatized by state-of-the-art AI-related technologies? Our work tackles this question through a data-driven approach: (i) developing a reproducible framework that exploits a battery of open-source Large Language Models to assess current AI and robotics' capabilities in performing job-related tasks; (ii) formalising and computing an AI exposure measure by occupation, namely the teai (Task Exposure to AI) index. Our results show that about one-third of U.S. employment is highly exposed to AI, primarily in high-skill jobs (aka, white collars). This exposure correlates positively with employment and wage growth from 2019 to 2023, indicating a beneficial impact of AI on productivity. The source codes and results are publicly available, enabling the whole community to benchmark and track AI and technology capabilities over time.
Open Source AI Has Founders--and the FTC--Buzzing
Y Combinator is famed for its Demo Days, where portfolio companies pitch their apps and wares in hopes of growing from a fledgling company into the next AirBnB. But on Thursday, the startup incubator hosted a mรฉlange of founders, venture capitalists, and US policy makers in its airy industrial space in San Francisco to tackle a defining topic for so many startups today: AI as the latest frontier in the battle between Big Tech and the little guys. For many early-stage tech entrepreneurs, questions around AI can carry existential weight. Ever since ChatGPT was unleashed in late 2022, OpenAI's technology, along with fast follows from Google's and Microsoft's AI teams, has dominated the conversation around this new era of artificial intelligence. But the increasing availability--and potency--of open source AI models has the potential to upend those dynamics.
X's Grok AI is scanning your tweets. Here's how to disable it
Elon Musk, the divisive CEO of Tesla and more recently the owner of Twitter (now known as X), is a fierce critic of the AI industry--but now also a deeply invested participant in that very same industry. X's Grok generative AI product is being integrated into the web and mobile versions of the social network, and training itself on billions of tweets thanks to an automatic opt-in for all users. Well, it seems like a constantly refreshed pool of conversations from some of the web's most active users was simply too much for company xAI to resist, which now automatically scans your "posts as well as your interactions, inputs, and [Grok search] results." At the moment, X is using Grok as a chatbot for premium users and to replace human-made summaries of late-breaking news stories, with predictable issues resulting. The flippant and "rebellious" tone of the Grok model's responses has been criticized by initial users, and its reliance on constantly updated data from X seems to make it particularly susceptible to deliberate misinformation campaigns.
Apple agrees to stick by Biden administration's voluntary AI safeguards
Apple has joined several other tech companies in agreeing to abide by voluntary AI safeguards laid out by the Biden administration. Those who make the pledge have committed to abide by eight guidelines related to safety, security and social responsibility, including flagging societal risks such as biases; testing for vulnerabilities, watermarking AI-generated images and audio; and sharing trust and safety details with the government and other companies. Amazon, Google, Microsoft and OpenAI were among the initial adoptees of the pact, which the White House announced last July. The voluntary agreement, which is not enforceable, will expire after Congress passes laws to regulate AI. Since the guidelines were announced, Apple unveiled a suite of AI-powered features under the umbrella name of Apple Intelligence.
Why Zuckerberg's multibillion-dollar gamble doesn't just matter to Meta
Spending on artificial intelligence could hit a staggering 1tn, according to analysts concerned about whether there will be a return on such a spree. Mark Zuckerberg's answer this week to such jitters was to release his latest AI system for free. Meta's Llama 3.1 405B is its most powerful yet, it says, and one of the most capable in the world. While the tech company didn't disclose how much it cost to train, Zuckerberg, its co-founder and chief executive, has previously disclosed a 10.5bn ( 8.9bn) investment in just the chips required to power its AI data centres โ with the rest of the electronics, the electricity itself, and the physical building an additional cost on top of that. Yet despite the exorbitant outlay, the parent of Facebook and Instagram will charge you nothing for it.
The Morning After: OpenAI reveals its AI-powered search engine, SearchGPT
OpenAI announced a new AI-powered search engine prototype called SearchGPT. It's described SearchGPT as "a temporary prototype of new AI search features that give you fast and timely answers with clear and relevant sources." The company plans to test out the product with 10,000 initial users, then roll it into ChatGPT after gathering feedback. It's a spicy time to launch AI-powered search engines. Last month, Perplexity faced criticism for summarizing stories from Forbes and Wired without adequate attribution or backlinks to the publications.
OpenAI takes on Google: Microsoft-backed tech giant launches an AI search tool dubbed SearchGPT
Google executives may be fearing the worst once again as Microsoft-backed rival OpenAI launches a new AI-powered search tool. 'SearchGPT', which is being trialed as a prototype before a wider rollout, scours the web for live news and information just like Google Search. OpenAI says the new product is particularly useful for queries about current events, recent developments, or specific information that ChatGPT might not know. Social media users have noted the parallels with the world's biggest search engine, with one saying'Google Search is definitely in trouble'. Another said: 'Anyone who has been paying attention knows there will be a new king of search within 10 years.
Training AI requires more data than we have -- generating synthetic data could help solve this challenge
Amritha R Warrier & AI4Media / Better Images of AI / error cannot generate / Licenced by CC-BY 4.0 The rapid rise of generative artificial intelligence like OpenAI's GPT-4 has brought remarkable advancements, but it also presents significant risks. One of the most pressing issues is model collapse, a phenomenon where AI models trained on largely AI-generated content tend to degrade over time. This degradation occurs as AI models lose information about their true underlying data distribution, resulting in increasingly similar and less diverse outputs full of biases and errors. As the internet becomes flooded with real-time AI-generated content, the scarcity of new, human-generated or natural data further exacerbates this problem. Without a steady influx of diverse, high-quality data, AI systems risk becoming less accurate and reliable.