Goto

Collaborating Authors

 Government


Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval

arXiv.org Artificial Intelligence

Retrieval-Augmented Generation (RAG) has become a popular technique for enhancing the reliability and utility of Large Language Models (LLMs) by grounding responses in external documents. Traditional RAG systems rely on Optical Character Recognition (OCR) to first process scanned documents into text. However, even state-of-the-art OCRs can introduce errors, especially in degraded or complex documents. Recent vision-language approaches, such as ColPali, propose direct visual embedding of documents, eliminating the need for OCR. This study presents a systematic comparison between a vision-based RAG system (ColPali) and more traditional OCR-based pipelines utilizing Llama 3.2 (90B) and Nougat OCR across varying document qualities. Beyond conventional retrieval accuracy metrics, we introduce a semantic answer evaluation benchmark to assess end-to-end question-answering performance. Our findings indicate that while vision-based RAG performs well on documents it has been fine-tuned on, OCR-based RAG is better able to generalize to unseen documents of varying quality. We highlight the key trade-offs between computational efficiency and semantic accuracy, offering practical guidance for RAG practitioners in selecting between OCR-dependent and vision-based document retrieval systems in production environments.


scDrugMap: Benchmarking Large Foundation Models for Drug Response Prediction

arXiv.org Artificial Intelligence

Drug resistance presents a major challenge in cancer therapy. Single cell profiling offers insights into cellular heterogeneity, yet the application of large-scale foundation models for predicting drug response in single cell data remains underexplored. To address this, we developed scDrugMap, an integrated framework featuring both a Python command-line interface and a web server for drug response prediction. scDrugMap evaluates a wide range of foundation models, including eight single-cell models and two large language models, using a curated dataset of over 326,000 cells in the primary collection and 18,800 cells in the validation set, spanning 36 datasets and diverse tissue and cancer types. We benchmarked model performance under pooled-data and cross-data evaluation settings, employing both layer freezing and Low-Rank Adaptation (LoRA) fine-tuning strategies. In the pooled-data scenario, scFoundation achieved the best performance, with mean F1 scores of 0.971 (layer freezing) and 0.947 (fine-tuning), outperforming the lowest-performing model by over 50%. In the cross-data setting, UCE excelled post fine-tuning (mean F1: 0.774), while scGPT led in zero-shot learning (mean F1: 0.858). Overall, scDrugMap provides the first large-scale benchmark of foundation models for drug response prediction in single-cell data and serves as a user-friendly, flexible platform for advancing drug discovery and translational research.


Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods

arXiv.org Artificial Intelligence

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for estimating model capabilities, they often fail to establish true upper bounds or predict deployment behavior. This literature review consolidates the rapidly evolving field of AI safety evaluations, proposing a systematic taxonomy around three dimensions: what properties we measure, how we measure them, and how these measurements integrate into frameworks. We show how evaluations go beyond benchmarks by measuring what models can do when pushed to the limit (capabilities), the behavioral tendencies exhibited by default (propensities), and whether our safety measures remain effective even when faced with subversive adversarial AI (control). These properties are measured through behavioral techniques like scaffolding, red teaming and supervised fine-tuning, alongside internal techniques such as representation analysis and mechanistic interpretability. We provide deeper explanations of some safety-critical capabilities like cybersecurity exploitation, deception, autonomous replication, and situational awareness, alongside concerning propensities like power-seeking and scheming. The review explores how these evaluation methods integrate into governance frameworks to translate results into concrete development decisions. We also highlight challenges to safety evaluations - proving absence of capabilities, potential model sandbagging, and incentives for "safetywashing" - while identifying promising research directions. By synthesizing scattered resources, this literature review aims to provide a central reference point for understanding AI safety evaluations.


GenAI in Entrepreneurship: a systematic review of generative artificial intelligence in entrepreneurship research: current issues and future directions

arXiv.org Artificial Intelligence

Generative Artificial Intelligence (GenAI) and Large Language Models (LLMs) are recognized to have significant effects on industry and business dynamics, not least because of their impact on the preconditions for entrepreneurship. There is still a lack of knowledge of GenAI as a theme in entrepreneurship research. This paper presents a systematic literature review aimed at identifying and analyzing the evolving landscape of research on the effects of GenAI on entrepreneurship. We analyze 83 peer-reviewed articles obtained from leading academic databases: Web of Science and Scopus. Using natural language processing and unsupervised machine learning techniques with TF-IDF vectorization, Principal Component Analysis (PCA), and hierarchical clustering, five major thematic clusters are identified: (1) Digital Transformation and Behavioral Models, (2) GenAI-Enhanced Education and Learning Systems, (3) Sustainable Innovation and Strategic AI Impact, (4) Business Models and Market Trends, and (5) Data-Driven Technological Trends in Entrepreneurship. Based on the review, we discuss future research directions, gaps in the current literature, as well as ethical concerns raised in the literature. We highlight the need for more macro-level research on GenAI and LLMs as external enablers for entrepreneurship and for research on effective regulatory frameworks that facilitate business experimentation, innovation, and further technology development.


Fact-checking Trump's claim of securing 10 trillion in investments for US

Al Jazeera

Since returning to the White House, US President Donald Trump has touted corporate and foreign US investment announcements as proof he is ushering in "the golden age of America". On January 21, Trump said that before he'd finished the "first full business day" of his second term, the United States had "already secured nearly 3 trillion of new investments". On April 2, he said, "It looks like we're going to have about 6 trillion of investments". Six days later, Trump told National Republican Congressional Committee Dinner attendees that the investment total was "now revised up to about 7 (trillion)". During an April 30 NewsNation town hall, Trump speculated that "it could be more than 8 trillion".


Army ditches helicopters for new radical air assault planes

FOX News

Fox News contributor Brett Velicovich joins'Fox & Friends First' to discuss Secretary's Hegseth's sweeping Army transformation, how Russia has responded to the U.S. minerals deal with Ukraine and the military bolstering drone technology. This is how the Army will island hop in the Pacific to fend off China. And by the way, Chinese President Xi Jinping has nothing like it. With a stunning announcement, the Army did more than ax 40 generals and open the door to AI. The Army bet its future on this radical aircraft, whose engines swivel to take off and land like a helicopter, or fly high and fast like an airplane.


US Marine Corps creates attack drone team as arms race with Russia, China heats up

FOX News

Fox News contributor and Army veteran Brett Velicovich shares insight into the United States' drone capabilities and how it compares to adversaries like Russia and China. The U.S. Marine Corps established an attack drone team earlier this year to respond to the rapid development of armed first-person view (FPV) drone technology and tactics, offering a glimpse into the evolving landscape of modern warfare and how future battles could be fought. The Marine Corps Attack Drone Team (MCADT) will be based at the Weapons Training Battalion, Marine Corps Base in Quantico, Virginia. The FPV drones used will offer squad-level lethality at a range of up to 20 kilometers, nearly 12.5 miles, for under 5,000, compared to more expensive weapons systems with less capability, according to a press release from the service. "MCADT is committed to rapidly integrating armed first-person view drones into the FMF [Fleet Marine Force], enhancing small-unit lethality and providing organic capabilities that warfighters currently lack," said Maj. Alejandro Tavizon, the headquarters company commander at Weapons Training Battalion and officer in charge of MCADT.


AI firms warned to calculate threat of super intelligence or risk it escaping human control

The Guardian

Artificial intelligence companies have been urged to replicate the safety calculations that underpinned Robert Oppenheimer's first nuclear test before they release all-powerful systems. Max Tegmark, a leading voice in AI safety, said he had carried out calculations akin to those of the US physicist Arthur Compton before the Trinity test and had found a 90% probability that a highly advanced AI would pose an existential threat. The US government went ahead with Trinity in 1945, after being reassured there was a vanishingly small chance of an atomic bomb igniting the atmosphere and endangering humanity. In a paper published by Tegmark and three of his students at the Massachusetts Institute of Technology (MIT), they recommend calculating the "Compton constant" – defined in the paper as the probability that an all-powerful AI escapes human control. In a 1959 interview with the US writer Pearl Buck, Compton said he had approved the test after calculating the odds of a runaway fusion reaction to be "slightly less" than one in three million.


ICE's Deportation Airline Hack Reveals Man 'Disappeared' to El Salvador

WIRED

A United States Customs and Border Protection request for information this week revealed the agency's plans to find vendors that can supply face recognition technology for capturing data on everyone entering the US in a vehicle like a car or van, not just the people sitting in the front seat. And a CBP spokesperson later told WIRED that the agency also has plans to expand its real-time face recognition capabilities at the border to detect people exiting the US as well--a focus that may be tied to the Trump administration's push to get undocumented people to "self-deport" and leave the US. WIRED also shed light this week on a recent CBP memo that rescinded a number of internal policies designed to protect vulnerable people--including pregnant women, infants, the elderly, and people with serious medical conditions--while in the agency's custody. Signed by acting commissioner Pete Flores, the order eliminates four Biden-era policies. Meanwhile, as the ripple effects of "SignalGate" continue, the communication app TeleMessage suspended "all services" pending an investigation after former US national security adviser Mike Waltz inadvertently called attention to the app, which subsequently suffered data breaches in recent days.


Rice-sized robot could make brain surgery safer and less invasive

FOX News

Surgeries may become safer and more precise than ever before. A French startup named Robeauté has just raised about 29 million to develop a truly groundbreaking neurosurgical microrobot. Imagine a device no bigger than a grain of rice that can carefully navigate the complex and delicate pathways of the brain. This little robot could change the way doctors treat brain tumors and other neurological conditions, making surgeries safer and more precise than ever before. Join The FREE CyberGuy Report: Get my expert tech tips, critical security alerts, and exclusive deals -- plus instant access to my free Ultimate Scam Survival Guide when you sign up! Brain surgery is incredibly complex.