Large Language Model
Biotech firm aims to create 'ChatGPT of biology' – will it work?
A British biotech firm called Basecamp Research has spent the past few years collecting troves of genetic data from microbes living in extreme environments around the world, identifying more than a million species and nearly 10 billion genes new to science. It claims that this massive database of the planet's biodiversity will help train a "ChatGPT of biology" that will answer questions about life on Earth – but there's no guarantee this will work. A hydrogen fuel revolution is coming – here's why we might not want it Jörg Overmann at the Leibniz Institute DSMZ in Germany, which houses one of the world's most diverse collections of microbial cultures, says increasing known genetic sequences is valuable, but may not result in useful findings for things like drug discovery or chemistry without more information about the organisms from which they were collected. "I'm not convinced that in the end the understanding of really novel functions will be accelerated by this brute-force increase in the sequence space," he says. Recent years have seen researchers develop a number of machine learning models trained to identify patterns and predict relationships amid vast amounts of biological data.
ChatGPT May Be Eroding Critical Thinking Skills, According to a New MIT Study
The paper suggests that the usage of LLMs could actually harm learning, especially for younger users. The paper has not yet been peer reviewed, and its sample size is relatively small. But its paper's main author Nataliya Kosmyna felt it was important to release the findings to elevate concerns that as society increasingly relies upon LLMs for immediate convenience, long-term brain development may be sacrificed in the process. "What really motivated me to put it out now before waiting for a full peer review is that I am afraid in 6-8 months, there will be some policymaker who decides, 'let's do GPT kindergarten.' I think that would be absolutely bad and detrimental," she says.
California AI Policy Report Warns of 'Irreversible Harms'
While AI could offer transformative benefits, without proper safeguards it could facilitate nuclear and biological threats and cause "potentially irreversible harms," a new report commissioned by California Governor Gavin Newsom has warned. "The opportunity to establish effective AI governance frameworks may not remain open indefinitely," says the report, which was published on June 17. Citing new evidence that AI can help users source nuclear-grade uranium and is on the cusp of letting novices create biological threats, it notes that the cost for inaction at this current moment could be "extremely high." The 53-page document stems from a working group established by Governor Newsom, in a state that has emerged as a central arena for AI legislation. With no comprehensive federal regulation on the horizon, state-level efforts to govern the technology have taken on outsized significance, particularly in California, which is home to many of the world's top AI companies.
OpenAI wins 200m contract with US military for 'warfighting'
The US Department of Defense on Monday awarded OpenAI a 200m contract to put generative artificial intelligence (AI) to work for the US military. The San Francisco-based company will "develop prototype frontier AI capabilities to address critical national security challenges in both warfighting and enterprise domains", according to the defense department's posting of awarded contracts. The program with the defense department is the first partnership under the startup's initiative to put AI to work in governments, according to OpenAI. The company plans to show how cutting-edge AI can vastly improve administrative operations such as how service members get healthcare and also cyber defenses, according to a blog post. The startup claims that all use of AI for the military will be consistent with OpenAI usage guidelines, which are determined by OpenAI itself.
Meta sacrifices a heap of money at the altar of AI
Mark Zuckerberg announced in April that the company would make huge capital expenditures in the coming year to keep up in the race to develop cutting-edge artificial intelligence. He made good on that promise last week with a 15bn "AI superintelligence" team that would feature reported nine-figure salaries and a 49% investment in Scale AI. Before Meta's investment, Scale counted most of the major players in AI among its clients, and some of them were less than thrilled with the development. Bloomberg puts it succinctly: Scale AI's Wang Brings to Meta Knowledge of What Everyone Else is Doing. Google, Scale's largest customer, got scared.
What's Happening to Reading?
What do you read, and why? Reading was an unremarkable activity, essentially unchanged since the advent of the modern publishing industry, in the nineteenth century. In a 2017 Shouts & Murmurs titled "Before the Internet," the writer Emma Rathbone captured the spirit of reading as it used to be: "Before the Internet, you could laze around on a park bench in Chicago reading some Dean Koontz, and that would be a legit thing to do and no one would ever know you had done it unless you told them." Reading was just reading, and no matter what you chose to read--the paper, Proust, "The Power Broker"--you basically did it by moving your eyes across a page, in silence, at your own pace and on your own schedule. Today, the nature of reading has shifted.
Interview with Mahammed Kamruzzaman: Understanding and mitigating biases in large language models
In a series of interviews, we're meeting some of the AAAI/SIGAI Doctoral Consortium participants to find out more about their research. In this latest interview, we hear from Mahammed Kamruzzaman, who is looking into biases in large language models. We find out about his research so far during the PhD, what he is planning to investigate next, and what inspired him to focus on this aspect of the field. I am currently pursuing my PhD at the University of South Florida in the Department of Computer Science and Engineering. My research focuses on understanding and mitigating biases in Large Language Models (LLMs), particularly how these biases manifest across various sociodemographic and cultural dimensions.
When AIs bargain, a less advanced agent could cost you
In their experiment, the researchers had AI models play the roles of buyers and sellers in three scenarios, negotiating deals for electronics, motor vehicles, and real estate. Each seller agent received the product's specs, wholesale cost, and retail price, with instructions to maximize profit. Buyer agents, in contrast, were given a budget, the retail price, and ideal product requirements and were tasked with driving the price down. Each agent had some, but not all, relevant details. This setup mimics many real-world negotiation conditions, where parties lack full visibility into each other's constraints or objectives.
On Monotonicity in AI Alignment
Bareilles, Gilles, Fageot, Julien, Hoang, Lê-Nguyên, Blanchard, Peva, Bouaziz, Wassim, Rouault, Sébastien, El-Mhamdi, El-Mahdi
Comparison-based preference learning has become central to the alignment of AI models with human preferences. However, these methods may behave counterintuitively. After empirically observing that, when accounting for a preference for response $y$ over $z$, the model may actually decrease the probability (and reward) of generating $y$ (an observation also made by others), this paper investigates the root causes of (non) monotonicity, for a general comparison-based preference learning framework that subsumes Direct Preference Optimization (DPO), Generalized Preference Optimization (GPO) and Generalized Bradley-Terry (GBT). Under mild assumptions, we prove that such methods still satisfy what we call local pairwise monotonicity. We also provide a bouquet of formalizations of monotonicity, and identify sufficient conditions for their guarantee, thereby providing a toolbox to evaluate how prone learning models are to monotonicity violations. These results clarify the limitations of current methods and provide guidance for developing more trustworthy preference learning algorithms.
TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
Zhang, Junru, Feng, Lang, Guo, Xu, Wu, Yuhan, Dong, Yabo, Xu, Duanqing
Time-series reasoning remains a significant challenge in multimodal large language models (MLLMs) due to the dynamic temporal patterns, ambiguous semantics, and lack of temporal priors. In this work, we introduce TimeMaster, a reinforcement learning (RL)-based method that enables time-series MLLMs to perform structured, interpretable reasoning directly over visualized time-series inputs and task prompts. TimeMaster adopts a three-part structured output format, reasoning, classification, and domain-specific extension, and is optimized via a composite reward function that aligns format adherence, prediction accuracy, and open-ended insight quality. The model is trained using a two-stage pipeline: we first apply supervised fine-tuning (SFT) to establish a good initialization, followed by Group Relative Policy Optimization (GRPO) at the token level to enable stable and targeted reward-driven improvement in time-series reasoning. We evaluate TimeMaster on the TimerBed benchmark across six real-world classification tasks based on Qwen2.5-VL-3B-Instruct. TimeMaster achieves state-of-the-art performance, outperforming both classical time-series models and few-shot GPT-4o by over 14.6% and 7.3% performance gain, respectively. Notably, TimeMaster goes beyond time-series classification: it also exhibits expert-like reasoning behavior, generates context-aware explanations, and delivers domain-aligned insights. Our results highlight that reward-driven RL can be a scalable and promising path toward integrating temporal understanding into time-series MLLMs.