Goto

Collaborating Authors

 FDA


This medical startup uses LLMs to run appointments and make diagnoses

MIT Technology Review

"Our focus is really on what we can do to pull the doctor out of the visit," says Akido's CTO. Imagine this: You've been feeling unwell, so you call up your doctor's office to make an appointment. At the appointment, you aren't rushed through describing your health concerns; instead, you have a full half hour to share your symptoms and worries and the exhaustive details of your health history with someone who listens attentively and asks thoughtful follow-up questions. You leave with a diagnosis, a treatment plan, and the sense that, for once, you've been able to discuss your health with the care that it merits. AI companies have stopped warning you that their chatbots aren't doctors Once cautious, OpenAI, Grok, and others will now dive into giving unverified medical advice with virtually no disclaimers. You might not have spoken to a doctor, or other licensed medical practitioner, at all.


Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy

arXiv.org Artificial Intelligence

Diabetic retinopathy (DR) is a leading cause of blindness worldwide, and AI systems can expand access to fundus photography screening. Current FDA-cleared systems primarily provide binary referral outputs, where this minimal output may limit clinical trust and utility. Yet, determining the most effective output format to enhance clinician-AI performance is an empirical challenge that is difficult to assess at scale. We evaluated multimodal large language models (MLLMs) for DR detection and their ability to simulate clinical AI assistance across different output types. Two models were tested on IDRiD and Messidor-2: GPT-4o, a general-purpose MLLM, and MedGemma, an open-source medical model. Experiments included: (1) baseline evaluation, (2) simulated AI assistance with synthetic predictions, and (3) actual AI-to-AI collaboration where GPT-4o incorporated MedGemma outputs. MedGemma outperformed GPT-4o at baseline, achieving higher sensitivity and AUROC, while GPT-4o showed near-perfect specificity but low sensitivity. Both models adjusted predictions based on simulated AI inputs, but GPT-4o's performance collapsed with incorrect ones, whereas MedGemma remained more stable. In actual collaboration, GPT-4o achieved strong results when guided by MedGemma's descriptive outputs, even without direct image access (AUROC up to 0.96). These findings suggest MLLMs may improve DR screening pipelines and serve as scalable simulators for studying clinical AI assistance across varying output configurations. Open, lightweight models such as MedGemma may be especially valuable in low-resource settings, while descriptive outputs could enhance explainability and clinician trust in clinical workflows.


Drug Repurposing Using Deep Embedded Clustering and Graph Neural Networks

arXiv.org Artificial Intelligence

Drug repurposing has historically been an economically infeasible process for identifying novel uses for abandoned drugs. Modern machine learning has enabled the identification of complex biochemical intricacies in candidate drugs; however, many studies rely on simplified datasets with known drug-disease similarities. We propose a machine learning pipeline that uses unsupervised deep embedded clustering, combined with supervised graph neural network link prediction to identify new drug-disease links from multi-omic data. Unsupervised autoencoder and cluster training reduced the dimensionality of omic data into a compressed latent embedding. A total of 9,022 unique drugs were partitioned into 35 clusters with a mean silhouette score of 0.8550. Graph neural networks achieved strong statistical performance, with a prediction accuracy of 0.901, receiver operating characteristic area under the curve of 0.960, and F1-Score of 0.901. A ranked list comprised of 477 per-cluster link probabilities exceeding 99 percent was generated. This study could provide new drug-disease link prospects across unrelated disease domains, while advancing the understanding of machine learning in drug repurposing studies.


Human-AI Collaboration Increases Efficiency in Regulatory Writing

arXiv.org Artificial Intelligence

Background: Investigational New Drug (IND) application preparation is time-intensive and expertise-dependent, slowing early clinical development. Objective: To evaluate whether a large language model (LLM) platform (AutoIND) can reduce first-draft composition time while maintaining document quality in regulatory submissions. Methods: Drafting times for IND nonclinical written summaries (eCTD modules 2.6.2, 2.6.4, 2.6.6) generated by AutoIND were directly recorded. For comparison, manual drafting times for IND summaries previously cleared by the U.S. FDA were estimated from the experience of regulatory writers ($\geq$6 years) and used as industry-standard benchmarks. Quality was assessed by a blinded regulatory writing assessor using seven pre-specified categories: correctness, completeness, conciseness, consistency, clarity, redundancy, and emphasis. Each sub-criterion was scored 0-3 and normalized to a percentage. A critical regulatory error was defined as any misrepresentation or omission likely to alter regulatory interpretation (e.g., incorrect NOAEL, omission of mandatory GLP dose-formulation analysis). Results: AutoIND reduced initial drafting time by $\sim$97% (from $\sim$100 h to 3.7 h for 18,870 pages/61 reports in IND-1; and to 2.6 h for 11,425 pages/58 reports in IND-2). Quality scores were 69.6\% and 77.9\% for IND-1 and IND-2. No critical regulatory errors were detected, but deficiencies in emphasis, conciseness, and clarity were noted. Conclusions: AutoIND can dramatically accelerate IND drafting, but expert regulatory writers remain essential to mature outputs to submission-ready quality. Systematic deficiencies identified provide a roadmap for targeted model improvements.


A Survey of Graph Neural Networks for Drug Discovery: Recent Developments and Challenges

arXiv.org Artificial Intelligence

Graph Neural Networks (GNNs) have gained traction in the complex domain of drug discovery because of their ability to process graph-structured data such as drug molecule models. This approach has resulted in a myriad of methods and models in published literature across several categories of drug discovery research. This paper covers the research categories comprehensively with recent papers, namely molecular property prediction, including drug-target binding affinity prediction, drug-drug interaction study, microbiome interaction prediction, drug repositioning, retrosynthesis, and new drug design, and provides guidance for future work on GNNs for drug discovery.


Evaluation of Machine Learning Reconstruction Techniques for Accelerated Brain MRI Scans

arXiv.org Artificial Intelligence

Figure 3: Distribution of Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Haar wavelet-based Perceptual Similarity Index (HaarPSI) scores for DeepFoqus-Accelerate reconstructions: (a-c) show results across 408 samples at 2x, 3x, and 4x acceleration, and (d-f) present distributions for the 36 image sets evaluated by reviewers. Figure 4: (A-B) Representative standard-of-care (SOC) images (first row) and DeepFoqus-Accelerate reconstructions from accelerated scans (second row), with corresponding quantitative and qualitative scores presented in the third row. Panel (B) shows two slices of the worst-case scenario in the qualitative dataset, characterized by wrap-around and motion artifacts. Discussion This evaluation of DeepFoqus-Accelerate demonstrates that this FDA-cleared k-space-based DL reconstruction software can reliably enable up to fourfold accelerated brain MRI acquisition without compromising diagnostic image quality. Both expert review and quantitative image similarity metrics confirm that AI-reconstructed images are clinically equivalent to fully sampled standards.


Moderna CEO Responds to RFK Jr.'s Crusade Against the Covid-19 Vaccine

WIRED

Speaking at a WIRED event Tuesday, Moderna CEO Stรฉphane Bancel said he was "encouraged" by the company's dialogue with the FDA--but acknowledged recent setbacks. Moderna CEO Stรฉphane Bancel prepares to testify before the Senate on March 22, 2023 in Washington, DC. At the WIRED Health summit on Tuesday, Moderna CEO Stรฉphane Bancel said the recent changes to Covid-19 vaccine policy made by Health and Human Services secretary Robert F. Kennedy, Jr. are a "step backward." Moderna is one of the manufacturers of mRNA-based Covid-19 vaccines, and last month the company received approval from the Food and Drug Administration for an updated version of the shot . But as part of that approval, the FDA imposed new restrictions on who can receive the vaccine.


Dangerous heart conditions detected in seconds with AI stethoscope

FOX News

Board-certified cardiothoracic surgeon Dr. Jeremy London, based in Savannah, Georgia, explains why VO2 max and muscle mass are the main indicators of longevity. The first artificial intelligence (AI) stethoscope has gone beyond listening to a heartbeat. Researchers at Imperial College London and Imperial College Healthcare NHS Trust discovered that an AI stethoscope can detect heart failure at an early stage. The TRICORDER study results, published in BMJ Journals, found that the AI-enabled stethoscope can help doctors identify three heart conditions in just 15 seconds. According to the British Heart Foundation (BHF), which partially funded the study, the researchers analyzed data from more than 1.5 million patients, focusing on people with heart failure symptoms like breathlessness, swelling and fatigue.


Resilient Biosecurity in the Era of AI-Enabled Bioweapons

arXiv.org Artificial Intelligence

Recent advances in generative biology have enabled the design of novel proteins, creating significant opportunities for drug discovery while also introducing new risks, including the potential development of synthetic bioweapons. Existing biosafety measures primarily rely on inference-time filters such as sequence alignment and protein-protein interaction (PPI) prediction to detect dangerous outputs. In this study, we evaluate the performance of three leading PPI prediction tools: AlphaFold 3, AF3Complex, and SpatialPPIv2. These models were tested on well-characterized viral-host interactions, such as those involving Hepatitis B and SARS-CoV-2. Despite being trained on many of the same viruses, the models fail to detect a substantial number of known interactions. Strikingly, none of the tools successfully identify any of the four experimentally validated SARS-CoV-2 mutants with confirmed binding. These findings suggest that current predictive filters are inadequate for reliably flagging even known biological threats and are even more unlikely to detect novel ones. We argue for a shift toward response-oriented infrastructure, including rapid experimental validation, adaptable biomanufacturing, and regulatory frameworks capable of operating at the speed of AI-driven developments.


Towards Early Detection: AI-Based Five-Year Forecasting of Breast Cancer Risk Using Digital Breast Tomosynthesis Imaging

arXiv.org Artificial Intelligence

As early detection of breast cancer strongly favors successful therapeutic outcomes, there is major commercial interest in optimizing breast cancer screening. However, current risk prediction models achieve modest performance and do not incorporate digital breast tomosynthesis (DBT) imaging, which was FDA-approved for breast cancer screening in 2011. To address this unmet need, we present a deep learning (DL)-based framework capable of forecasting an individual patient's 5-year breast cancer risk directly from screening DBT. Using an unparalleled dataset of 161,753 DBT examinations from 50,590 patients, we trained a risk predictor based on features extracted using the Meta AI DINOv2 image encoder, combined with a cumulative hazard layer, to assess a patient's likelihood of developing breast cancer over five years. On a held-out test set, our best-performing model achieved an AUROC of 0.80 on predictions within 5 years. These findings reveal the high potential of DBT-based DL approaches to complement traditional risk assessment tools, and serve as a promising basis for additional investigation to validate and enhance our work.