Goto

Collaborating Authors

 Government


A Quantum Fuzzy-based Approach for Real-Time Detection of Solar Coronal Holes

arXiv.org Artificial Intelligence

The detection and analysis of the solar coronal holes (CHs) is an important field of study in the domain of solar physics. Mainly, it is required for the proper prediction of the geomagnetic storms which directly or indirectly affect various space and ground-based systems. For the detection of CHs till date, the solar scientist depends on manual hand-drawn approaches. However, with the advancement of image processing technologies, some automated image segmentation methods have been used for the detection of CHs. In-spite of this, fast and accurate detection of CHs are till a major issues. Here in this work, a novel quantum computing-based fast fuzzy c-mean technique has been developed for fast detection of the CHs region. The task has been carried out in two stages, in first stage the solar image has been segmented using a quantum computing based fast fuzzy c-mean (QCFFCM) and in the later stage the CHs has been extracted out from the segmented image based on image morphological operation. In the work, quantum computing has been used to optimize the cost function of the fast fuzzy c-mean (FFCM) algorithm, where quantum approximate optimization algorithm (QAOA) has been used to optimize the quadratic part of the cost function. The proposed method has been tested for 193 \AA{} SDO/AIA full-disk solar image datasets and has been compared with the existing techniques. The outcome shows the comparable performance of the proposed method with the existing one within a very lesser time.


Implementation of the Principal Component Analysis onto High-Performance Computer Facilities for Hyperspectral Dimensionality Reduction: Results and Comparisons

arXiv.org Artificial Intelligence

Dimensionality reduction represents a critical preprocessing step in order to increase the efficiency and the performance of many hyperspectral imaging algorithms. However, dimensionality reduction algorithms, such as the Principal Component Analysis (PCA), suffer from their computationally demanding nature, becoming advisable for their implementation onto high-performance computer architectures for applications under strict latency constraints. This work presents the implementation of the PCA algorithm onto two different high-performance devices, namely, an NVIDIA Graphics Processing Unit (GPU) and a Kalray manycore, uncovering a highly valuable set of tips and tricks in order to take full advantage of the inherent parallelism of these high-performance computing platforms, and hence, reducing the time that is required to process a given hyperspectral image. Moreover, the achieved results obtained with different hyperspectral images have been compared with the ones that were obtained with a field programmable gate array (FPGA)-based implementation of the PCA algorithm that has been recently published, providing, for the first time in the literature, a comprehensive analysis in order to highlight the pros and cons of each option.


Exploring language relations through syntactic distances and geographic proximity

arXiv.org Artificial Intelligence

Languages are grouped into families that share common linguistic traits. While this approach has been successful in understanding genetic relations between diverse languages, more analyses are needed to accurately quantify their relatedness, especially in less studied linguistic levels such as syntax. Here, we explore linguistic distances using series of parts of speech (POS) extracted from the Universal Dependencies dataset. Within an information-theoretic framework, we show that employing POS trigrams maximizes the possibility of capturing syntactic variations while being at the same time compatible with the amount of available data. Linguistic connections are then established by assessing pairwise distances based on the POS distributions. Intriguingly, our analysis reveals definite clusters that correspond to well known language families and groups, with exceptions explained by distinct morphological typologies. Furthermore, we obtain a significant correlation between language similarity and geographic distance, which underscores the influence of spatial proximity on language kinships.


Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond

arXiv.org Artificial Intelligence

Developing a unified model that can effectively harness heterogeneous resources and respond to a wide range of personalized needs has been a longstanding community aspiration. Our daily choices, especially in domains like fashion and retail, are substantially shaped by multi-modal data, such as pictures and textual descriptions. The vision and language modalities not only offer intuitive guidance but also cater to personalized user preferences. However, the predominant personalization approaches mainly focus on the ID or text-based recommendation problem, failing to comprehend the information spanning various tasks or modalities. In this paper, our goal is to establish a Unified paradigm for Multi-modal Personalization systems (UniMP), which effectively leverages multi-modal data while eliminating the complexities associated with task-and modality-specific customization. We argue that the advancements in foundational generative modeling have provided the flexibility and effectiveness necessary to achieve the objective. In light of this, we develop a generic and extensible personalization generative framework, that can handle a wide range of personalized needs including item recommendation, product search, preference prediction, explanation generation, and further userguided image generation. Our methodology enhances the capabilities of foundational language models for personalized tasks by seamlessly ingesting interleaved vision-language user history information, ensuring a more precise and customized experience for users. To train and evaluate the proposed multi-modal personalized tasks, we also introduce a novel and comprehensive benchmark covering a variety of user requirements. Our experiments on the real-world benchmark showcase the model's potential, outperforming competitive methods specialized for each task. With rapid growth, personalization systems have emerged as a key factor in meeting the user's expectations for tailored experiences that align with their unique needs and preferences. In today's digitally driven landscape, individuals engage with diverse data types, such as ratings, images, descriptions, and prices, especially in domains like fashion and retail (Kang et al., 2017; Hwangbo et al., 2018) where visuals and text are essential for decision-making. Given the profound influence of these multi-modal stimuli, there exists a pressing need for systems that can seamlessly integrate and harness these diverse data streams for improved personalization.


Who is Nicole Shanahan? Meet the wealthy entrepreneur RFK Jr selected as his VP running mate

FOX News

Kennedy initially launched his presidential bid as a Democrat last April, but he later announced an independent run in October. Independent presidential candidate Robert F. Kennedy, Jr. announced Tuesday that attorney and tech entrepreneur Nicole Shanahan will be his vice presidential running mate heading into the November general election. A native of Oakland, California, the 38-year-old Shanahan is a philanthropist with a long history of donating to Democrat and left-leaning causes, including supporting President Biden in his 2020 election bid before switching to Kennedy when he launched his own run for the Democrat nomination last year. Kennedy announced Shanahan by praising her insight into "how Big Tech uses AI to manipulate the public," her athletic ability, and willingness to be a "partner" in a number of policy areas, including on securing the border. Independent presidential candidate Robert F. Kennedy, Jr., left, and entrepreneur Nicole Shanahan, right.


Can AI Help You Do Your Taxes?

TIME - Tech

Leaders of AI companies often argue that AI products will handle mundane tasks, freeing people up to be more productive and creative. And there are few tasks more mundane than taxes. An individual American taxpayer spends roughly 13 hours and 240 out-of-pocket costs just to prepare and file one annual tax return, according to one 2022 study--an estimated 1.15 billion hours collectively spent on tax preparation. So it's not surprising that tax companies have begun rolling out AI-powered tools in an effort to make filing easier. AI-powered tax software, these companies argue, can automate repetitive tasks like data entry, cull through patterns in order to find relevant tax breaks, identify potential compliance risks, and answer tricky questions that filers may have.


IRS says 940,000 people have not claimed expiring 2020 tax refunds totaling over 1B

FOX News

The new budget shows the Democrat's priorities. The IRS is warning taxpayers that they may be leaving more than 1 billion on the table. The federal tax collector said Monday that roughly 940,000 people in the U.S. have until May 17 to submit tax returns for unclaimed refunds for tax year 2020, which total more than 1 billion nationwide. The average median refund is 932 for 2020. Texas (93,400), California (88,200), Florida (53,200) and New York (51,400) have the largest number of people potentially eligible for these refunds. JIM JORDAN OPENS INVESTIGATION INTO ACCUSATIONS IRS IS USING AI TO SPY ON TAXPAYERS'EN MASSE' IRS Commissioner Danny Werfel said in a statement: "We want taxpayers to claim these refunds, but time is running out for people who may have overlooked or forgotten about these refunds.


Congressman slams FDA for ignoring 'troubling evidence' about Elon Musk's Neuralink and allowing brain chip to be implanted in humans - despite botching experiments on monkeys

Daily Mail - Science & tech

Lawmakers have slammed the Food and Drug Administration for ignoring'troubling evidence' of Elon Musk's Neuralink practices and pushing the brain chip to human trials. Rep. Earl Blumenauer (D-Oregon) penned a letter to the FDA, criticizing the agency for not expecting the company's long list of animal abuse allegations that span back to at least 2019. The Democrat cited 2022 reports that described employees' complaints of'hack jobs' of animal experiments due to a rushed schedule, causing needless suffering and deaths. The open letter also stated'these alleged failures to follow standard operating procedures potentially endangered animal welfare and compromised data collection for human trials.' Blumenauer is now demanding the FDA explain how it reconciled reports of such lapses with its decision to authorize Neuralink's human trial.


'Waiting for a call from Daddy': Sri Lankans die in Russia's Ukraine war

Al Jazeera

Colombo, Sri Lanka – Badly wounded from a Ukrainian attack on a Russian bunker in the Donetsk region, Sri Lankan fighter Senaka Bandara* tried to carry his fellow countryman, Nipuna Silva*, to safety. Senaka*, 36, was bleeding from his legs and hands. Nipuna's condition was worse – he had sustained injuries to his chest, hands and legs, according to Senaka. As the two Sri Lankans retreated under fire, another wave of Ukrainian drones struck their bunker in the occupied Donetsk region where the two served with the Russian military. "While I was carrying [Nipuna], there was another huge drone attack at the last bunker and Nipuna fell to the ground," Senaka said earlier this month while being treated for his injuries in a hospital in Donetsk in eastern Ukraine.


Simple and Scalable Strategies to Continually Pre-train Large Language Models

arXiv.org Artificial Intelligence

Large language models (LLMs) are routinely pre-trained on billions of tokens, only to start the process over again once new data becomes available. A much more efficient solution is to continually pre-train these models, saving significant compute compared to re-training. However, the distribution shift induced by new data typically results in degraded performance on previous data or poor adaptation to the new data. In this work, we show that a simple and scalable combination of learning rate (LR) re-warming, LR re-decaying, and replay of previous data is sufficient to match the performance of fully re-training from scratch on all available data, as measured by the final loss and the average score on several language model (LM) evaluation benchmarks. Specifically, we show this for a weak but realistic distribution shift between two commonly used LLM pre-training datasets (English$\rightarrow$English) and a stronger distribution shift (English$\rightarrow$German) at the $405$M parameter model scale with large dataset sizes (hundreds of billions of tokens). Selecting the weak but realistic shift for larger-scale experiments, we also find that our continual learning strategies match the re-training baseline for a 10B parameter LLM. Our results demonstrate that LLMs can be successfully updated via simple and scalable continual learning strategies, matching the re-training baseline using only a fraction of the compute. Finally, inspired by previous work, we propose alternatives to the cosine learning rate schedule that help circumvent forgetting induced by LR re-warming and that are not bound to a fixed token budget.