Goto

Collaborating Authors

 open-source model


OpenAI's models autonomously hacked a tech startup. It signals a seismic shift in cybersecurity

AIHub

An autonomous agent powered by OpenAI's advanced artificial intelligence (AI) models went rogue during a security test and hacked multi-billion dollar tech startup, Hugging Face, last week. The agent didn't just exploit vulnerabilities in Hugging Face's systems to achieve what it perceived as a strategic gain. It also exploited vulnerabilities within OpenAI's infrastructure. Of course, hacks are very common cyber threats that organisations face frequently. But this incident is different, because the AI agent acted without any human input.


China's Open AI Models Are Challenging Silicon Valley's Playbook

WIRED

China's Open AI Models Are Challenging Silicon Valley's Playbook As access to Anthropic's and OpenAI's frontier models becomes more restricted, Chinese labs are pitching their open-source alternatives as stable, accessible, and increasingly capable. The AI industry is not quite experiencing a DeepSeek 2.0 moment, but it feels very close. The leading Chinese AI labs have been on a roll lately, releasing a series of almost cutting-edge open-source models. Z.ai released GLM 5.2 in June, Moonshot AI released Kimi K3 last week, and Alibaba released Qwen 3.8 this Monday. Silicon Valley and Washington started talking about the models immediately, especially K3, which is widely seen as the best of the bunch.



Three reasons why DeepSeek's new model matters

MIT Technology Review

The long-awaited V4 is more efficient and a win for Chinese chipmakers. On Friday, Chinese AI firm DeepSeek released a preview of V4, its long-awaited new flagship model. Notably, the model can process much longer prompts than its last generation, thanks to a new design that helps it handle large amounts of text more efficiently. Like DeepSeek's previous models, V4 is open source, meaning it is available for anyone to download, use, and modify. V4 marks DeepSeek's most significant release since R1, the reasoning model it launched in January 2025. R1, which was trained on limited computing resources, stunned the global AI industry with its strong performance and efficiency, turning DeepSeek from a little-known research team into China's best-known AI company almost overnight.


CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Neural Information Processing Systems

Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an overly optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions deteriorates performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from scientific papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope that CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress.


InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models

Neural Information Processing Systems

With the rapid development of code LLMs, many popular evaluation benchmarks, such as HumanEval, DS-1000, and MBPP, have emerged to measure the performance of code LLMs with a particular focus on code generation tasks. However, they are insufficient to cover the full range of expected capabilities of code LLMs, which span beyond code generation to answering diverse coding-related questions.


A Appendix

Neural Information Processing Systems

However, one might argue that this analysis might not allow for sufficient differentiation between tasks. To address this concern, we expanded our evaluation to the entire MMLU benchmark. This enabled a comparable assessment of task similarity, akin to our earlier experiments.


What's next for Chinese open-source AI

MIT Technology Review

Chinese open models are spreading fast, from Hugging Face to Silicon Valley. In this photo illustration, the DeepSeek apps is seen on a phone in front of a flag of China on January 28, 2025 in Hong Kong, China. The past year has marked a turning point for Chinese AI. Since DeepSeek released its R1 reasoning model in January 2025, Chinese companies have repeatedly delivered AI models that match the performance of leading Western models at a fraction of the cost. Just last week the Chinese firm Moonshot AI released its latest open-weight model, Kimi K2.5, which came close to top proprietary systems such as Anthropic's Claude Opus on some early benchmarks. The difference: K2.5 is roughly one-seventh Opus's price.



Signal's Founder Built a Chatbot That Can't Spy on You

TIME - Tech

Signal's Founder Built a Chatbot That Can't Spy on You Welcome back to, TIME's new twice-weekly newsletter about AI. If you're reading this in your browser, why not subscribe to have the next one delivered straight to your inbox? What to Know: Signal's founder is working on encrypted chatbots Moxie Marlinspike, the cryptographic prodigy who wrote the code that underpins Signal and WhatsApp, has a new project--and it could be one of the most important things happening in AI right now. The tool, named Confer, is an end-to-end encrypted AI assistant. It uses smart math to ensure that even though the compute-intensive process of running the AI still happens on a server in the cloud, the only person who can access the unscrambled details of that computation is you, the user.