Goto

Collaborating Authors

 Large Language Model


Detection of LLM-Generated Java Code Using Discretized Nested Bigrams

arXiv.org Artificial Intelligence

Large Language Models (LLMs) are currently used extensively to generate code by professionals and students, motivating the development of tools to detect LLM-generated code for applications such as academic integrity and cybersecurity. We address this authorship attribution problem as a binary classification task along with feature identification and extraction. We propose new Discretized Nested Bigram Frequency features on source code groups of various sizes. Compared to prior work, improvements are obtained by representing sparse information in dense membership bins. Experimental evaluation demonstrated that our approach significantly outperformed a commonly used GPT code-detection API and baseline features, with accuracy exceeding 96% compared to 72% and 79% respectively in detecting GPT-rewritten Java code fragments for 976 files with GPT 3.5 and GPT4 using 12 features. We also outperformed three prior works on code author identification in a 40-author dataset. Our approach scales well to larger data sets, and we achieved 99% accuracy and 0.999 AUC for 76,089 files and over 1,000 authors with GPT 4o using 227 features.


A Layered Multi-Expert Framework for Long-Context Mental Health Assessments

arXiv.org Artificial Intelligence

Long-form mental health assessments pose unique challenges for large language models (LLMs), which often exhibit hallucinations or inconsistent reasoning when handling extended, domain-specific contexts. We introduce Stacked Multi-Model Reasoning (SMMR), a layered framework that leverages multiple LLMs and specialized smaller models as coequal 'experts'. Early layers isolate short, discrete subtasks, while later layers integrate and refine these partial outputs through more advanced long-context models. We evaluate SMMR on the DAIC-WOZ depression-screening dataset and 48 curated case studies with psychiatric diagnoses, demonstrating consistent improvements over single-model baselines in terms of accuracy, F1-score, and PHQ-8 error reduction. By harnessing diverse 'second opinions', SMMR mitigates hallucinations, captures subtle clinical nuances, and enhances reliability in high-stakes mental health assessments. Our findings underscore the value of multi-expert frameworks for more trustworthy AI-driven screening.


ChatGPT's new AI search beats Google in this one thing

PCWorld

OpenAI's ChatGPT has removed the last barrier to using ChatGPT as a search engine, the requirement to log in. OpenAI launched the feature last fall, but required a login. Now, the feature can be used without the need for registration. ChatGPT Search is essentially just ChatGPT, and can be accessed at ChatGPT.com. But below the "Message ChatGPT" box you'll see a small icon called "Search" that can be clicked.


OpenAI co-founder John Schulman has left Anthropic after less than a year

Engadget

Less than a year into his tenure at the company, OpenAI co-founder John Schulman is leaving Anthropic. The startup confirmed Schulman's departure after The Information, Reuters and other publications reported on the exit. "We are sad to see John go but fully support his decision to pursue new opportunities and wish him all the very best," said Jared Kaplan, Anthropic's chief science officer, in a statement the company shared with Engadget. Schulman left OpenAI last August alongside Peter Deng, the company's former vice-president of consumer product. Schulman is considered one of the original architects of ChatGPT.


DeepSeek limits model access due to overwhelming server demand

Engadget

DeepSeek recent explosion in popularity continues to be a problem for the AI startup. In a notification spotted by Bloomberg, the company said it was temporarily limiting access to its application programming interface service in response to a shortage of server capacity. "Due to current server resource constraints, we have temporarily suspended API service recharges to prevent any potential impact on your operations," DeepSeek said. "Existing balances can still be used for calls. Separately, DeepSeek announced pricing for its chat model would increase to 0.27 per million input tokens and 1.10 per million output tokens starting February 8. DeepSeek has been dealing with overwhelming demand for its services since the debut of its R1 model on January 20.


Lyft uses Anthropic's Claude chatbot to handle user complaints

Engadget

Lyft is partnering with Anthropic to bring the startup's AI tech to its platform. "Anthropic, known for its human-centric approach to AI, will work with Lyft to build smart, safe, and empathetic AI-powered products that put riders and drivers first," the two said in a joint press release. If you're a frequent Lyft rider, you can see the early results of that collaboration when you go through the company's customer care AI assistant, which features integration with Anthrophic's Claude chatbot. According to Lyft, the tool is already helping to resolve thousands of customer issues every day, and has reduced average resolution times by 87 percent. Moving forward, Lyft plans to integrate Anthropic's tech across its business.


Which countries have banned DeepSeek and why?

Al Jazeera

This week, government agencies in countries including South Korea and Australia have blocked access to Chinese artificial intelligence (AI) startup DeepSeek's new AI chatbot programme, mostly for government employees. Other countries, including the United States, have said they may also seek to block DeepSeek from government employees' mobile devices, according to media reports. All cite "security concerns" about the Chinese technology and a lack of clarity about how users' personal information is handled by the operator. Last month, DeepSeek made headlines after it caused share prices in US tech companies to plummet, after it claimed that its model would cost only a fraction of the money its competitors had spent on their own AI programmes to build. The news caused social media users to joke: "I can't believe ChatGPT lost its job to AI." Here's what we know about DeepSeek and why countries are banning it.


Reframing digital transformation through the lens of generative AI

MIT Technology Review

Enterprise adoption of generative AI technologies has undergone explosive growth in the last two years and counting. Powerful solutions underpinned by this new generation of large language models (LLMs) have been used to accelerate research, automate content creation, and replace clunky chatbots with AI assistants and more sophisticated AI agents that closely mimic human interaction. "In 2023 and the first part of 2024, we saw enterprises experimenting, trying out new use cases to see, 'What can this new technology do for me?'" explains Arthy Krishnamurthy, senior director for business transformation at Dataiku. But while many organizations were eager to adopt and exploit these exciting new capabilities, some may have underestimated the need to thoroughly scrutinize AI-related risks and recalibrate existing frameworks and forecasts for digital transformation.


ASUS's Zenfone 12 Ultra leans heavily into AI

Engadget

The Zenfone 12 Ultra, announced today, is ASUS's latest flagship smartphone, and much like its competitors, it leans hard into AI. Thanks to a Snapdragon 8 Elite, the Zenfone 12 Ultra can perform AI tasks offline and online through the cloud, including transcribing audio, summarizing articles and documents and providing real-time interpretation on calls for supported languages. It can also use Circle to Search much like other Android phones. The onboard AI is powered by Meta's Llama 3 8B language model, which works without an internet connection. The Zenfone 12 Ultra's FHD AMOLED display measures 6.78 inches and has a standard refresh rate of up to 120Hz under normal operation, and up to 144Hz while gaming.


South Korean ministries and police block access to DeepSeek

The Japan Times

South Korean ministries and police said Thursday they were blocking access to DeepSeek on work computers after the Chinese artificial intelligence startup did not respond to a data watchdog request about how it manages user information. DeepSeek launched its R1 chatbot last month, claiming it matches the capacity of AI pacesetters in the United States at a fraction of the investment, upending the global industry. South Korea, along with countries such as France and Italy, have asked questions about DeepSeek's data practices, submitting a written request for information about how the company handles user information.