Large Language Model
Watch out, Nvidia. OpenAI's proprietary AI chip is coming along
According to a new report from Reuters (spotted by Thurrott), OpenAI could finalize the design of its first 3nm AI chip in the coming months, with the goal of starting mass production at TSMC in 2026. The chip is being developed by a team of 40 OpenAI employees in collaboration with Broadcom. The project is being led by Richard Ho, OpenAI's new head of hardware, who previously worked on solutions for Google's infrastructure and cloud services. According to Reuters, OpenAI's chip will be able to both train and run AI models, but initially it'll be used mainly for inference (running AI models) and to a limited extent within the company's infrastructure. Demand for Nvidia's AI chips remains extremely high right now, with companies like OpenAI, Microsoft, Meta, and Google investing billions in AI data centers.
Fox News AI Newsletter: VP calls for ideology-free AI
Gladstone A.I. co-founders and CEOs Edouard Harris and Jeremie Harris explain the major role that A.I will play in national security and warfare on'The Will Cain Show.' Vice President JD Vance will attend an AI summit in Paris, France, a French official said anonymously. FREE FROM BIAS: Vice President JD Vance told world leaders in Paris on Tuesday that the United States intends to remain the dominant force in artificial intelligence and warned that the European Union's far tougher regulatory approach to the technology could cripple it. 'TRYING TO SLOW US DOWN': OpenAI CEO Sam Altman said Elon Musk is "probably just trying to slow us down" with his bid to purchase the company, insisting on Tuesday that it is not for sale. 'MASS SURVEILLANCE': OpenAI CEO Sam Altman predicts that artificial general intelligence will lead to lower costs for many goods, but has also warned that AI could be leveraged by authoritarian governments aiming to control people. TRANSLATED TRUTH: Whether you have an iPhone or an Android, these apps have got you covered with features like live speech translation, text input and even AI-powered sign and menu translation.
Sam Altman applauds JD Vance's AI speech in Paris, illustrates ways to take advantage of 'remarkable' tech
Vice President JD Vance addressed the AI Action Summit in Paris Tuesday during his first foreign trip since taking office. OpenAI CEO Sam Altman commended Vice President JD Vance's artificial intelligence (AI) speech in Paris on Tuesday while laying out his vision for how people can take advantage of the rapidly evolving technology at the same conference. Altman and Vance appeared Tuesday at the AI Action Summit in Paris, where world leaders, top tech executives and policymakers teamed up to hash out tech policy and its intersection with global security, economics and governance. During his remarks, Vance called for AI systems developed in the U.S. to remain free of "ideological bias" and vowed that the U.S. would "never restrict our citizens' right to free speech." Vance also pushed for a "deregulatory flavor" to emerge at the conference while cautioning against the pitfalls of "excessive regulation" that could hamper a transformative industry.
Want to run AI on your PC? You're gonna need a bigger hard drive
When people talk about the "size" of an AI model, they're referring to the number of "parameters" it contains. A parameter is one variable in the AI model that determines how it generates output, and any given AI model can have billions of these parameters. Also referred to as model weights, these parameters occupy storage space to operate properly -- and when an AI model has billions of parameters, storage requirements can quickly balloon. As you can see, the storage space consumed by an LLM increases with the size of its parameters. The same is true for other types of generative AI models, too.
Elon Musk's A.I.-Fuelled War on Human Agency
Not long ago, the American public could have been forgiven for thinking of Elon Musk's vaunted Department of Government Efficiency (DOGE) as a version of a familiar Republican cost-cutting, government-shrinking project. The man who took over Twitter and slashed its staff by around eighty per cent would take a similarly aggressive tack against bureaucratic inefficiency, reining in budgets and laying off federal employees. In the past couple of weeks, though, it's become clear that Musk's aim within the Trump Administration goes further: he wants not only to reduce the U.S. government but to install his own technological vision of the future at its heart. To run his agency, Musk brought on a group of tech-company managers and inexperienced twentysomethings whose credentials included internships at SpaceX. We watched as this crew began interrogating federal employees about their jobs, interfering with the system that controls payments at the Treasury Department, and trawling government budgets while Musk used X, the social platform he owns, to call out the agencies and programs in his crosshairs.
Elon Musk owning OpenAI would be a terrible idea. That doesn't mean it won't happen Chris Stokel-Walker
The two had a blowout argument over the future direction of OpenAI โ the company they came together to found in 2015 โ with Altman seemingly content to pursue a for-profit approach and Musk feeling that was forswearing the founding principles of the firm as well as its name. OpenAI couldn't be open, he reckoned, if it was closed off and trying to make money rather than better humanity. So it's no surprise that Musk, who lodged an audacious bid to take over Twitter a little more than two years ago, which ended up with his ownership of the platform now called X, has sought to put a spoiler in two years of near-untrammelled growth for OpenAI. Musk โ who is currently overhauling (to his supporters; "tearing down" to his opponents) the US government to be, as he would describe it, leaner and more efficient while also devastating important programmes such as international aid and cutting-edge scientific research โ has lodged a near 100bn bid for OpenAI's non-profit arm. "It's time for OpenAI to return to the open-source, safety-focused force for good it once was," Musk said in a statement supplied by the lawyer shepherding his bid.
AI feud: How Musk and Altman's partnership turned toxic
The feud between Elon Musk and Sam Altman has become one of the bitterest rivalries in business history, with the Tesla tycoon bidding to buy Altman's OpenAI in an apparent attempt to derail the ChatGPT maker's ascent to becoming one of the world's most important companies. Musk and Altman were among the 11-person team that founded OpenAI in 2015. Created as a counterweight to Google's dominance in artificial intelligence, the project got its initial funding from Musk, who invested 45 million to get it started. Three years later, Musk departed OpenAI. The company initially cited "a potential future conflict for Elon ... as Tesla continues to become more focused on AI," noting the electric vehicle company's ambitions in autonomous driving.
Spectral Journey: How Transformers Predict the Shortest Path
Cohen, Andrew, Gromov, Andrey, Yang, Kaiyu, Tian, Yuandong
Decoder-only transformers lead to a step-change in capability of large language models. However, opinions are mixed as to whether they are really planning or reasoning. A path to making progress in this direction is to study the model's behavior in a setting with carefully controlled data. Then interpret the learned representations and reverse-engineer the computation performed internally. We study decoder-only transformer language models trained from scratch to predict shortest paths on simple, connected and undirected graphs. In this setting, the representations and the dynamics learned by the model are interpretable. We present three major results: (1) Two-layer decoder-only language models can learn to predict shortest paths on simple, connected graphs containing up to 10 nodes. (2) Models learn a graph embedding that is correlated with the spectral decomposition of the line graph. (3) Following the insights, we discover a novel approximate path-finding algorithm Spectral Line Navigator (SLN) that finds shortest path by greedily selecting nodes in the space of spectral embedding of the line graph.
The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
Yu, Mo, Liu, Lemao, Wu, Junjie, Chung, Tsz Ting, Zhang, Shunchi, Li, Jiangnan, Yeung, Dit-Yan, Zhou, Jie
In a systematic way, we investigate a widely asked question: Do LLMs really understand what they say?, which relates to the more familiar term Stochastic Parrot. To this end, we propose a summative assessment over a carefully designed physical concept understanding task, PhysiCo. Our task alleviates the memorization issue via the usage of grid-format inputs that abstractly describe physical phenomena. The grids represents varying levels of understanding, from the core phenomenon, application examples to analogies to other abstract patterns in the grid world. A comprehensive study on our task demonstrates: (1) state-of-the-art LLMs, including GPT-4o, o1 and Gemini 2.0 flash thinking, lag behind humans by ~40%; (2) the stochastic parrot phenomenon is present in LLMs, as they fail on our grid task but can describe and recognize the same concepts well in natural language; (3) our task challenges the LLMs due to intrinsic difficulties rather than the unfamiliar grid format, as in-context learning and fine-tuning on same formatted data added little to their performance.
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
Xu, Jiacheng, Pang, Bo, Qu, Jin, Hayashi, Hiroaki, Xiong, Caiming, Zhou, Yingbo
Software testing is a critical aspect of software development, yet generating test cases remains a routine task for engineers. This paper presents a benchmark, CLOVER, to evaluate models' capabilities in generating and completing test cases under specific conditions. Spanning from simple assertion completions to writing test cases that cover specific code blocks across multiple files, these tasks are based on 12 python repositories, analyzing 845 problems with context lengths ranging from 4k to 128k tokens. Utilizing code testing frameworks, we propose a method to construct retrieval contexts using coverage information. While models exhibit comparable performance with short contexts, notable differences emerge with 16k contexts. Notably, models like GPT-4o and Claude 3.5 can effectively leverage relevant snippets; however, all models score below 35\% on the complex Task III, even with the oracle context provided, underscoring the benchmark's significance and the potential for model improvement. The benchmark is containerized for code execution across tasks, and we will release the code, data, and construction methodologies.