Generative AI
From Human Hands to Robotic Limbs: A Study in Motor Skill Embodiment for Telemanipulation
Shi, Haoyi, Su, Mingxi, Morris, Ted, Morellas, Vassilios, Papanikolopoulos, Nikolaos
Abstract-- This paper presents a teleoperation system for controlling a redundant degree-of-freedom (DOF) robot manipulator using human arm gestures. We propose a GRU-based Variational Autoencoder (VAE) to learn a latent representation of the manipulator's configuration space, capturing its complex joint kinematics. A fully-connected neural network maps human arm configurations into this latent space, allowing the system to mimic and generate corresponding manipulator trajectories in real-time through the VAE decoder. Arrow shows the mapping relationship between the manipulator's For example, an operator can use as agriculture, healthcare medicine, warehousing, and manufacturing. Another have proliferated, such as for image generation and natural approach instead uses an external RGB and RGBD (depth) language generation, for example ChatGPT, Midjourney, camera to estimate the operator's 6-DOF hand pose [4], [5] and Dall-E, to name a few.
e-SimFT: Alignment of Generative Models with Simulation Feedback for Pareto-Front Design Exploration
Cheong, Hyunmin, Ataei, Mohammadmehdi, Khasahmadi, Amir Hosein, Jayaraman, Pradeep Kumar
Deep generative models have recently shown success in solving complex engineering design problems where models predict solutions that address the design requirements specified as input. However, there remains a challenge in aligning such models for effective design exploration. For many design problems, finding a solution that meets all the requirements is infeasible. In such a case, engineers prefer to obtain a set of Pareto optimal solutions with respect to those requirements, but uniform sampling of generative models may not yield a useful Pareto front. To address this gap, we introduce a new framework for Pareto-front design exploration with simulation fine-tuned generative models. First, the framework adopts preference alignment methods developed for Large Language Models (LLMs) and showcases the first application in fine-tuning a generative model for engineering design. The important distinction here is that we use a simulator instead of humans to provide accurate and scalable feedback. Next, we propose epsilon-sampling, inspired by the epsilon-constraint method used for Pareto-front generation with classical optimization algorithms, to construct a high-quality Pareto front with the fine-tuned models. Our framework, named e-SimFT, is shown to produce better-quality Pareto fronts than existing multi-objective alignment methods.
Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription
Alkaeed, Mahdi, Abioye, Sofiat, Qayyum, Adnan, Mekki, Yosra Magdi, Berrou, Ilhem, Abdallah, Mohamad, Al-Fuqaha, Ala, Bilal, Muhammad, Qadir, Junaid
In response to the success of proprietary Large Language Models (LLMs) such as OpenAI's GPT-4, there is a growing interest in developing open, non-proprietary LLMs and AI foundation models (AIFMs) for transparent use in academic, scientific, and non-commercial applications. Despite their inability to match the refined functionalities of their proprietary counterparts, open models hold immense potential to revolutionize healthcare applications. In this paper, we examine the prospects of open-source LLMs and AIFMs for developing healthcare applications and make two key contributions. Firstly, we present a comprehensive survey of the current state-of-the-art open-source healthcare LLMs and AIFMs and introduce a taxonomy of these open AIFMs, categorizing their utility across various healthcare tasks. Secondly, to evaluate the general-purpose applications of open LLMs in healthcare, we present a case study on personalized prescriptions. This task is particularly significant due to its critical role in delivering tailored, patient-specific medications that can greatly improve treatment outcomes. In addition, we compare the performance of open-source models with proprietary models in settings with and without Retrieval-Augmented Generation (RAG). Our findings suggest that, although less refined, open LLMs can achieve performance comparable to proprietary models when paired with grounding techniques such as RAG. Furthermore, to highlight the clinical significance of LLMs-empowered personalized prescriptions, we perform subjective assessment through an expert clinician. We also elaborate on ethical considerations and potential risks associated with the misuse of powerful LLMs and AIFMs, highlighting the need for a cautious and responsible implementation in healthcare.
Microsoft's latest AI feature may just stop working. Here's why
Microsoft is indeed making access to OpenAI's 01 AI reasoning model completely free -- but there's a limitation and one which Microsoft is refusing to tell you about. Microsoft said Wednesday that it would provide access to OpenAI's o1 model, for free, to Copilot users as part of a toggle option called "Think Deeper." OpenAI uses the o1 model in its paid ChatGPT plans and charges 20/mo for "limited" access to the model, and unlimited access to it for a Pro plan costing 200/mo. However, Mustafa Suleyman, the chief of Microsoft's AI, said that the model would be "free and available," and "everywhere at no cost" -- potentially an enormous discount. Microsoft said Friday that there are limits to their new Think Deeper feature, which they're keeping mum about.
The Download: following DeepSeek's lead, and OpenAI's new research agent
When the Chinese firm DeepSeek dropped a large language model called R1 two weeks ago, it sent shock waves through the US tech industry. Not only did R1 match the best of the homegrown competition, it was built for a fraction of the cost--and given away for free. DeepSeek has now suddenly become the company to beat. What exactly did it do to rattle the tech world so fully? And what can we learn from the buzz about what's coming next?
OpenAI's new agent can compile detailed reports on practically any topic
OpenAI claims the tool represents a significant step toward its overarching goal of developing artificial general intelligence (AGI) that matches (or surpasses) human performance. It says that what takes the tool "tens of minutes" would take a human many hours. In response to a single query, such as "Draw me up a competitive analysis between streaming platforms," Deep Research will search the web, analyze the information it encounters, and compile a detailed report that cites its sources. It's also able to draw from files uploaded by users. OpenAI developed Deep Research using the same "chain of thought" reinforcement-learning methods it used to create its o1 multistep reasoning model. But while o1 was designed to focus primarily on mathematics, coding, or other STEM-based tasks, Deep Research can tackle a far broader range of subjects.
SoftBank forms joint venture with OpenAI in enterprise play
SoftBank Group will spend 3 billion a year to adopt and deploy OpenAI technology throughout its operations, while the two companies have agreed to form a joint venture to market the artificial intelligence as an enterprise solution. "This initiative will not only transform the way SoftBank Group operates but also revolutionize the way companies work in Japan and around the globe," SoftBank CEO Masayoshi Son said in a statement Monday. The technology, which the company describes as an advanced enterprise AI called Cristal intelligence, will be used at all companies under the SoftBank group, including Arm, Line and PayPay, to improve productivity and drive innovation. For instance, SoftBank's telecom unit plans to make more than 100 million workflows automated, the company said in the press release.
ChatGPT's Deep Research tool can create reports from hundreds of online sources
Two days after releasing o3-mini to the world, the company made a surprise announcement on Sunday evening, revealing Deep Research. The new feature allows ChatGPT to find, analyze and synthesize hundreds of websites and online sources to create reports "at the level of a research analyst." The chatbot will then take "anywhere from 5 to 30 minutes" to compile an answer, a side panel documenting the agent's progress and citations as it works. "It accomplishes in tens of minutes what would take a human many hours," OpenAI says of the new feature. "Our ultimate aspiration is a model that can uncover and discover new knowledge for itself," said Mark Chen, chief research officer at OpenAI, during the company's reveal livestream.
s1: Simple test-time scaling
Muennighoff, Niklas, Yang, Zitong, Shi, Weijia, Li, Xiang Lisa, Fei-Fei, Li, Hajishirzi, Hannaneh, Zettlemoyer, Luke, Liang, Percy, Candรจs, Emmanuel, Hashimoto, Tatsunori
Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts. We seek the simplest approach to achieve test-time scaling and strong reasoning performance. First, we curate a small dataset s1K of 1,000 questions paired with reasoning traces relying on three criteria we validate through ablations: difficulty, diversity, and quality. Second, we develop budget forcing to control test-time compute by forcefully terminating the model's thinking process or lengthening it by appending "Wait" multiple times to the model's generation when it tries to end. This can lead the model to double-check its answer, often fixing incorrect reasoning steps. After supervised finetuning the Qwen2.5-32B-Instruct language model on s1K and equipping it with budget forcing, our model s1-32B exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24). Further, scaling s1-32B with budget forcing allows extrapolating beyond its performance without test-time intervention: from 50% to 57% on AIME24. Our model, data, and code are open-source at https://github.com/simplescaling/s1
VILP: Imitation Learning with Latent Video Planning
Xu, Zhengtong, Qiu, Qiang, She, Yu
In the era of generative AI, integrating video generation models into robotics opens new possibilities for the general-purpose robot agent. This paper introduces imitation learning with latent video planning (VILP). We propose a latent video diffusion model to generate predictive robot videos that adhere to temporal consistency to a good degree. Our method is able to generate highly time-aligned videos from multiple views, which is crucial for robot policy learning. Our video generation model is highly time-efficient. For example, it can generate videos from two distinct perspectives, each consisting of six frames with a resolution of 96x160 pixels, at a rate of 5 Hz. In the experiments, we demonstrate that VILP outperforms the existing video generation robot policy across several metrics: training costs, inference speed, temporal consistency of generated videos, and the performance of the policy. We also compared our method with other imitation learning methods. Our findings indicate that VILP can rely less on extensive high-quality task-specific robot action data while still maintaining robust performance. In addition, VILP possesses robust capabilities in representing multi-modal action distributions. Our paper provides a practical example of how to effectively integrate video generation models into robot policies, potentially offering insights for related fields and directions. For more details, please refer to our open-source repository https://github.com/ZhengtongXu/VILP.