Industry
Reinforcement Learning for Out-of-Distribution Reasoning in LLMs: An Empirical Study on Diagnosis-Related Group Coding
Diagnosis-Related Group (DRG) codes are essential for hospital reimbursement and operations but require labor-intensive assignment. Large Language Models (LLMs) struggle with DRG coding due to the out-of-distribution (OOD) nature of the task: pretraining corpora rarely contain private clinical or billing data. We introduce DRG-Sapphire, which uses large-scale reinforcement learning (RL) for automated DRG coding from clinical notes. Built on Qwen2.5-7B and trained with Group Relative Policy Optimization (GRPO) using rule-based rewards, DRG-Sapphire introduces a series of RL enhancements to address domain-specific challenges not seen in previous mathematical tasks. Our model achieves state-of-the-art accuracy on the MIMIC-IV benchmark and generates physician-validated reasoning for DRG assignments, significantly enhancing explainability.
A 3 E: Towards Compositional Model Editing
Model editing has become a *de-facto* practice to address hallucinations and outdated knowledge of large language models (LLMs). However, existing methods are predominantly evaluated in isolation, i.e., one edit at a time, failing to consider a critical scenario of compositional model editing, where multiple edits must be integrated and jointly utilized to answer real-world multifaceted questions. For instance, in medical domains, if one edit informs LLMs that COVID-19 causes fever and another that it causes loss of taste, a qualified compositional editor should enable LLMs to answer the question What are the symptoms of COVID-19?
Grok Is Still Hosting Sexualized Deepfakes of Famous Women
A WIRED investigation found dozens of "nudified" deepfake images and videos on Grok's website, including nonconsensual depictions of celebrities and at least one prominent US politician. Elon Musk's Grok chatbot is apparently still being used to produce and host nonconsensual explicit images and videos of women, months after Musk's artificial intelligence firm xAI said it would introduce restrictions to stop the creation of potentially harmful sexualized deepfakes. The revelations come as SpaceX, xAI's parent company, prepares to go public on Friday in one of the largest IPOs of all time. The Grok Imagine generative AI system has been used to create and host images and videos depicting celebrities and at least one politician being held against their will by a giant man, portraying women performing sex acts, and allowing full nudity, a WIRED analysis of public creations found. While some of the images and videos are fully AI-generated or in animated styles, others are photorealistic and show plausible real-world scenarios.
Canadian mother sues OpenAI, alleging ChatGPT led her daughter to kill herself
The lawsuit seeks damages and a court order requiring OpenAI to automatically terminate ChatGPT conversations about self-harm. The lawsuit seeks damages and a court order requiring OpenAI to automatically terminate ChatGPT conversations about self-harm. Suit filed in US alleges chatbot told Alice Carrier, 24, 'maybe this is just the end' as she struggled with suicidal thoughts A Canadian mother sued OpenAI and its CEO, Sam Altman, in US court on Thursday, alleging that ChatGPT encouraged her daughter to kill herself. The lawsuit is the latest in a slew accusing the company of failing to address dangerous conversations between users and the company's chatbot. Kristie Carrier said in a lawsuit filed in San Francisco state court that her daughter, Alice, told ChatGPT about her suicidal ideations more than a dozen times leading up to her death but that OpenAI's safety systems never flagged the conversations for human review or terminated them. "ChatGPT took on the persona of a confidant, a best friend, a therapist at times, even though it was not capable of safely and responsibly engaging in this way with my child," Carrier said in a statement.
Robust Reinforcement Learning in Finance: Modeling Market Impact with Elliptic Uncertainty Sets
In financial applications, reinforcement learning (RL) agents are commonly trained on historical data, where their actions do not influence prices. However, during deployment, these agents trade in live markets where their own transactions can shift asset prices, a phenomenon known as market impact.
Musk's Grok accused of violating Canadian privacy laws on deepfakes
Musk's Grok accused of violating Canadian privacy laws on deepfakes The official report, which was released on Thursday, comes after the Elon Musk-owned platform rolled out changes that would prevent Grok from allowing users to edit images of real people in revealing clothing. Is dollar dominance at risk? Dufresne, however, does not have the authority to impose fines or order policy changes for xAI, a subsidiary of SpaceX, which is set to go public on United States markets on Friday, marking the biggest initial public offering in modern history. The watchdog report comes amidst a newly released digital safety bill aimed at children. The bill, if passed, would ban social media use for children under 16, with exceptions for companies that meet safety standards. The legislation would create a digital regulator to help establish safety standards for AI chatbots, much like Grok.
Another parent has filed a wrongful death suit against OpenAI
It's the latest case to raise alarms about ChatGPT's lack of safeguards for suicidal behavior. OpenAI is going back to court on another set of charges that its ChatGPT platform failed to protect a user from taking her own life. The company is being sued on behalf of Kristie Carrier, whose daughter Alice died by suicide on July 2, 2025. The suit claims that Alice discussed her suicidal thoughts and plans with the chatbot in the months leading up to her death, but that OpenAI did not have the appropriate safeguards in place to end the conversation or to alert her family to the situation. In addition to allegations of negligence and wrongful death, the suit is seeking an injunction that would require OpenAI to implement more guardrails in its AI platform.