Goto

Collaborating Authors

 Large Language Model


This 30 bundle helps you write, plan & create easier with ChatGPT

PCWorld

Artificial intelligence has changed the way we write, create, work, and plan--and with the 2025 ChatGPT Skills & Creativity Bundle, you'll be at the forefront of this revolution. For a limited time, you can access five expert-led courses that teach you how to supercharge your creativity, productivity, and AI skills for just 29.99 instead of 249. Whether you're an aspiring content creator, business professional, or someone who simply wants to make life easier with AI, this bundle will equip you with practical skills to get the most out of ChatGPT. From crafting engaging blog posts and video scripts to enhancing office productivity, bulk content creation, and even travel planning, this bundle offers real-world applications that will help you streamline your work and spark your creativity. The flagship course, Creative Writing & Content Creation with ChatGPT, guides you through advanced prompt engineering techniques to generate high-quality, engaging content while maintaining your unique voice.


South Korea removes DeepSeek from app stores pending privacy review

Al Jazeera

South Korea has suspended downloads of DeepSeek's artificial intelligence-powered chatbot pending a review of the Chinese start-up's privacy standards. South Korea's privacy watchdog said on Monday that DeepSeek's R1 chatbot was removed from the local versions of Apple's App Store and Google Play after the Hangzhou-based firm acknowledged that it had failed to comply with personal data protection rules. The Personal Information Protection Commission said in a statement that DeepSeek accepted its proposal to suspend downloads of the app. The chatbot is still available for those who have already downloaded the app. "To prevent further concerns from spreading, the commission recommended that DeepSeek temporarily suspend its service while making the necessary improvements," the commission said, adding that bringing the app in line with local regulations would "inevitably take a significant amount of time".


DeepSeek Not Available for Download in South Korea as Authorities Address Privacy Concerns

TIME - Tech

DeepSeek, a Chinese artificial intelligence startup, has temporarily paused downloads of its chatbot apps in South Korea while it works with local authorities to address privacy concerns, according to South Korean officials on Monday. South Korea's Personal Information Protection Commission said DeepSeek's apps were removed from the local versions of Apple's App Store and Google Play on Saturday evening and that the company agreed to work with the agency to strengthen privacy protections before relaunching the apps. Read More: Is the DeepSeek Panic Overblown? The action does not affect users who have already downloaded DeepSeek on their phones or use it on personal computers. Nam Seok, director of the South Korean commission's investigation division, advised South Korean users of DeepSeek to delete the app from their devices or avoid entering personal information into the tool until the issues are resolved.


Xi-Jack Ma chat seen as next catalyst for blistering China rally

The Japan Times

A potential encounter this week between Chinese President Xi Jinping and e-commerce icon Jack Ma, coming after a blistering run by tech shares, could be the next catalyst to extend the rally in China's stocks. Prominent entrepreneurs including Ma have been invited to meet the nation's top leaders, people familiar with the matter said last week. The potential show of support for the private sector coincides with the recent surge in equities in Hong Kong, driven by growing capabilities in artificial intelligence. The Hang Seng China Enterprises Index jumped 4.1% on Friday to its highest since February 2022, exceeding an October peak spurred by a stimulus blitz. A tech gauge in Hong Kong entered a bull market earlier this month, fueled by Chinese startup DeepSeek's AI model that's hailed as a game-changer.


Can Musk damage OpenAI even though his bid has failed?

BBC News

It was a huge sum - but less than the 157bn the firm was valued at in a funding round just four months ago, and much lower than the 300bn that some think it is worth now. Complicating all of this is OpenAI's unusual structure which involves a partnership between non-profit and for-profit arms. Mr Altman is understood to want to change that, stripping it of its non-profit board. That involves costs which Mr Musk is seemingly trying to inflate. "What Musk is trying to do here is raise the perceived value of the non-profit arm of OpenAI, so that OpenAI has to pay more to get out of the obligations it has to its own non-profit," said Dr Penn.


FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training

arXiv.org Artificial Intelligence

Selecting high-quality data can significantly improve the pretraining efficiency of large language models (LLMs). Existing methods generally rely on heuristic techniques and single-quality signals, limiting their ability to evaluate data quality comprehensively. In this work, we propose FIRE, a flexible and scalable framework for integrating multiple data quality raters, which allows for a comprehensive assessment of data quality across various dimensions. FIRE aligns multiple quality signals into a unified space, and integrates diverse data quality raters to provide a comprehensive quality signal for each data point. Further, we introduce a progressive data selection scheme based on FIRE that iteratively refines the selection of high-quality data points. Experiments on the SlimPajama dataset reveal that FIRE outperforms other data selection methods and significantly enhances the pretrained model across a wide range of downstream tasks, with a 2.9% average performance improvement over Random and reducing the FLOPs necessary to achieve a certain performance level by more than half.


Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study

arXiv.org Artificial Intelligence

Recently, inference-time scaling of chain-of-thought (CoT) has been demonstrated as a promising approach for addressing multi-modal reasoning tasks. While existing studies have predominantly centered on text-based thinking, the integration of both visual and textual modalities within the reasoning process remains unexplored. In this study, we pioneer the exploration of inference-time scaling with multi-modal thought, aiming to bridge this gap. To provide a comprehensive analysis, we systematically investigate popular sampling-based and tree search-based inference-time scaling methods on 10 challenging tasks spanning various domains. Besides, we uniformly adopt a consistency-enhanced verifier to ensure effective guidance for both methods across different thought paradigms. Results show that multi-modal thought promotes better performance against conventional text-only thought, and blending the two types of thought fosters more diverse thinking. Despite these advantages, multi-modal thoughts necessitate higher token consumption for processing richer visual inputs, which raises concerns in practical applications. We hope that our findings on the merits and drawbacks of this research line will inspire future works in the field.


QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models

arXiv.org Artificial Intelligence

Language Models (LLMs) are often quantized to lower precision to reduce the memory cost and latency in inference. However, quantization often degrades model performance, thus fine-tuning is required for various down-stream tasks. Traditional fine-tuning methods such as stochastic gradient descent and Adam optimization require backpropagation, which are error-prone in the low-precision settings. To overcome these limitations, we propose the Quantized Zeroth-Order (QuZO) framework, specifically designed for fine-tuning LLMs through low-precision (e.g., 4- or 8-bit) forward passes. Our method can avoid the error-prone low-precision straight-through estimator, and utilizes optimized stochastic rounding to mitigate the increased bias. QuZO simplifies the training process, while achieving results comparable to first-order methods in ${\rm FP}8$ and superior accuracy in ${\rm INT}8$ and ${\rm INT}4$ training. Experiments demonstrate that low-bit training QuZO achieves performance comparable to MeZO optimization on GLUE, Multi-Choice, and Generation tasks, while reducing memory cost by $2.94 \times$ in LLaMA2-7B fine-tuning compared to quantized first-order methods.


One for All: A General Framework of LLMs-based Multi-Criteria Decision Making on Human Expert Level

arXiv.org Artificial Intelligence

Multi-Criteria Decision Making~(MCDM) is widely applied in various fields, using quantitative and qualitative analyses of multiple levels and attributes to support decision makers in making scientific and rational decisions in complex scenarios. However, traditional MCDM methods face bottlenecks in high-dimensional problems. Given the fact that Large Language Models~(LLMs) achieve impressive performance in various complex tasks, but limited work evaluates LLMs in specific MCDM problems with the help of human domain experts, we further explore the capability of LLMs by proposing an LLM-based evaluation framework to automatically deal with general complex MCDM problems. Within the framework, we assess the performance of various typical open-source models, as well as commercial models such as Claude and ChatGPT, on 3 important applications, these models can only achieve around 60\% accuracy rate compared to the evaluation ground truth. Upon incorporation of Chain-of-Thought or few-shot prompting, the accuracy rates rise to around 70\%, and highly depend on the model. In order to further improve the performance, a LoRA-based fine-tuning technique is employed. The experimental results show that the accuracy rates for different applications improve significantly to around 95\%, and the performance difference is trivial between different models, indicating that LoRA-based fine-tuned LLMs exhibit significant and stable advantages in addressing MCDM tasks and can provide human-expert-level solutions to a wide range of MCDM challenges.


Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer

arXiv.org Artificial Intelligence

We propose Rotate, Clip, and Partition (RCP), a quantization-aware training (QAT) approach that first realizes extreme compression of LLMs with W2A4KV4(2-bit weight, 4-bit activation, and 4-bit KV cache) configuration. RCP integrates recent rotation techniques with a novel non-uniform weight quantizer design, by quantitatively analyzing the impact of random rotation on 2-bit weight quantization. Our weight quantizer features Learnable Direct Partitioning (LDP), which introduces learnable parameters to directly learn non-uniform intervals jointly with LLM weights. We also present a specialized GPU kernel that supports GEMV on non-uniform W2A4. Experiments show that RCP can compress LLaMA-2-7B to W2A4KV4 with a loss of only 2.84 WikiText2 ppl and 5.29 times reduced memory footprint. Furthermore, RCP can quantize challenging mobile-targeted LLaMA-3.2 models and domain-specific WizardCoder-7B and MetaMath-7B with no critical problems such as convergence failure and repetition. Code will be made available at blind_review.