Large Language Model
Training Large Language Models for Advanced Typosquatting Detection
Since the early days of the commercial internet, typosquatting has exploited the simplest of human errors, mistyping a URL, to serve as a potent tool for cybercriminals. Initially observed as an opportunistic tactic, typosquatting involves registering domain names that closely match that of reputable brands, thereby redirecting users to counterfeit websites. This has evolved into a sophisticated form of cyberattack used to conduct phishing schemes, distribute malware, and harvest sensitive data. Now with billions of domain names and TLDs in circulation, the scale and impact of typosquatting have grown exponentially. This poses significant risks to individuals, businesses, and national cybersecurity infrastructure. This whitepaper explores how emerging large language model (LLM) techniques can enhance the detection of typosquatting attempts, ultimately fortifying defenses against one of the internet's most enduring cyber threats. Cybercriminals employ various domain squatting techniques to deceive users and bypass traditional security measures. These methods include but not limited to: Character Substitution: These attacks swap similar looking characters like replacing "o" with "0" in go0gle[.]com to trick users into believing they are visiting the legitimate site. Omission or Addition: This method involves removing or adding a character, creating domains such as gogle[.]com
A Refined Analysis of Massive Activations in LLMs
Owen, Louis, Chowdhury, Nilabhra Roy, Kumar, Abhay, Gรผra, Fabian
Motivated in part by their relevance for low-precision training and quantization, massive activations in large language models (LLMs) have recently emerged as a topic of interest. However, existing analyses are limited in scope, and generalizability across architectures is unclear. This paper helps address some of these gaps by conducting an analysis of massive activations across a broad range of LLMs, including both GLU-based and non-GLU-based architectures. Our findings challenge several prior assumptions, most importantly: (1) not all massive activations are detrimental, i.e. suppressing them does not lead to an explosion of perplexity or a collapse in downstream task performance; (2) proposed mitigation strategies such as Attention KV bias are model-specific and ineffective in certain cases. We consequently investigate novel hybrid mitigation strategies; in particular pairing Target Variance Rescaling (TVR) with Attention KV bias or Dynamic Tanh (DyT) successfully balances the mitigation of massive activations with preserved downstream model performance in the scenarios we investigated. Our code is available at: https://github.com/bluorion-com/refine_massive_activations.
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
Karamolegkou, Antonia, Nikandrou, Malvina, Pantazopoulos, Georgios, Villegas, Danae Sanchez, Rust, Phillip, Dhar, Ruchira, Hershcovich, Daniel, Sรธgaard, Anders
This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such technologies. Despite a high adoption rate of these models, our findings highlight concerns related to contextual understanding, cultural sensitivity, and complex scene understanding, particularly for individuals who may rely solely on them for visual interpretation. Informed by these results, we collate five user-centred tasks with image and video inputs, including a novel task on Optical Braille Recognition. Our systematic evaluation of twelve MLLMs reveals that further advancements are necessary to overcome limitations related to cultural context, multilingual support, Braille reading comprehension, assistive object recognition, and hallucinations. This work provides critical insights into the future direction of multimodal AI for accessibility, underscoring the need for more inclusive, robust, and trustworthy visual assistance technologies.
How This Tool Could Decode AI's Inner Mysteries
The scientists didn't have high expectations when they asked their AI model to complete the poem. "He saw a carrot and had to grab it," they prompted the model. "His hunger was like a starving rabbit," it replied. The rhyming couplet wasn't going to win any poetry awards. But when the scientists at AI company Anthropic inspected the records of the model's neural network, they were surprised by what they found.
Anthropic can now track the bizarre inner workings of a large language model
It's no secret that large language models work in mysterious ways. Few--if any--mass-market technologies have ever been so little understood. That makes figuring out what makes them tick one of the biggest open challenges in science. Shedding some light on how these models work would expose their weaknesses, revealing why they make stuff up and can be tricked into going off the rails. It would help resolve deep disputes about exactly what these models can and can't do.
Fox News 'Antisemitism Exposed' Newsletter: Does 'AI' stand for 'anti-Israel'?
UPenn Wharton School Associate Professor Ethan Mollick weighs in on the Biden White House's new guidelines for artificial intelligence in the workplace on'Fox News Live.' Fox News' "Antisemitism Exposed" newsletter brings you stories on the rising anti-Jewish prejudice across the U.S. and the world. IN TODAY'S NEWSLETTER: - ADL issues'urgent call' alleging anti-Israel bias in 4 AI large language models - Georgetown grad student accused of spreading Hamas propaganda - Israeli hostages' families sue Mahmoud Khalil, Columbia organizers as alleged'Hamas' propaganda arm' The ADL's report found that virtually all artifical intelligence tools displayed a built-in bias against Israel and Jews. TOP STORY: A new report from the Anti-Defamation League (ADL) shows anti-Jewish and anti-Israel biases among AI large language models. The organization used thousands of AI queries to find "a concerning inability to accurately reject antisemitic tropes and conspiracy theories." Additionally, every LLM except GPT showed bias regarding Jewish conspiracy theories and even more bias against Israel than Jews, the ADL said.
The Importance of Distrust in Trusting Digital Worker Chatbots
Adopting and implementing digital automation technologies, including artificial intelligence (AI) models such as ChatGPT, robotic process automation (RPA), and other emerging AI technologies, will revolutionize many industries and business models. It is forecasted that the rise of AI will impact a wide range of job functions and roles. White-collar positions such as administrative, customer service, and back-office roles will all be impacted by AI-fueled digital automation. The adoption of digital workers is currently positioned in the early adopter phase of the product lifecycle.1 AI-driven digital workers are expected to substantially alter many white-collar tasks, including finance, customer support, human resources, sales, and marketing.42 A study from Oxford University and Deloitte predicts AI is a significant risk to the white-collar workforce.
OpenAI releases impressive 4o image generator for free and paid users
Earlier this week, OpenAI released their "most advanced image generator yet" and made it available through ChatGPT using the GPT-4o model. ChatGPT previously relied on Dall-E to generate images. According to OpenAI, the improved 4o model is able to produce precise, accurate, and photorealistic results. They claim that it's also particularly good at rendering text, following instructions precisely, and even understanding the context of a chat. All of this includes the transformation of uploaded images or using uploaded images as visual inspiration.
OpenAI delays rollout of ChatGPT's image generator to free users
Free ChatGPT users will have to wait a while longer to be able to use its built-in image generation capability. OpenAI has just launched a feature that will allow users to generate images directly inside of ChatGPT, and it was supposed to roll out to all Plus, Pro, Team and Free users. But according to company CEO Sam Altman, it has been way more popular than OpenAI had expected even though they already had high expectations to begin with. As such, its rollout to the free tier is "unfortunately going to be delayed for a while." People have been posting ChatGPT's output all over social media.
What is vibe coding, should you be doing it, and does it matter?
Getting an AI to write software for you? Want to write software, but haven't got the first clue where to start? Enter "vibe coding", a term that has swept the internet to describe the use of AI tools, including large language models (LLMs) like ChatGPT, to generate computer code even if you can't program. "Vibe coding basically refers to using generative AI not just to assist with coding, but to generate the entire code for an app," says Noah Giansiracusa at Bentley University in Waltham, Massachusetts. Users ask, or prompt, LLM-based models such as ChatGPT, Claude or Copilot to produce the code for an app or service, and the AI system does all the work.