Goto

Collaborating Authors

 Generative AI


Will AI help or hurt workers? At SXSW, that depends on who you ask.

The Japan Times

At the screenings of Ryan Gosling's new movie "The Fall Guy" and Sydney Sweeney's "Immaculate" -- headlining events of the first week of this year's South by Southwest (SXSW) conference -- a reel of speakers touting the merits of AI including Peter Deng, head of ChatGPT at OpenAI, was loudly booed by audiences. In other halls of the annual Austin confab, which attracts over 300,000 people each year, the tone was one of heady optimism. Companies were spinning a positive narrative -- pushing the idea that AI wouldn't destroy jobs, but rather represented a once-in-a-generation opportunity to modernize a multitude of industries. International Business Machines Chief Human Resources Officer Nickle LaMoreaux told a SXSW panel that AI will instead alleviate labor shortages that are set to worsen as birth rates decline across developed economies. And while only a tiny percentage of jobs can be completely automated, the vast majority won't disappear -- but will change dramatically.


BlendScape: Enabling Unified and Personalized Video-Conferencing Environments through Generative AI

arXiv.org Artificial Intelligence

Today's video-conferencing tools support a rich range of professional and social activities, but their generic, grid-based environments cannot be easily adapted to meet the varying needs of distributed collaborators. To enable end-user customization, we developed BlendScape, a system for meeting participants to compose video-conferencing environments tailored to their collaboration context by leveraging AI image generation techniques. BlendScape supports flexible representations of task spaces by blending users' physical or virtual backgrounds into unified environments and implements multimodal interaction techniques to steer the generation. Through an evaluation with 15 end-users, we investigated their customization preferences for work and social scenarios. Participants could rapidly express their design intentions with BlendScape and envisioned using the system to structure collaboration in future meetings, but experienced challenges with preventing distracting elements. We implement scenarios to demonstrate BlendScape's expressiveness in supporting distributed collaboration techniques from prior work and propose composition techniques to improve the quality of environments.


Apple's MM1 AI Model Shows a Sleeping Giant Is Waking Up

WIRED

While the tech industry went gaga for generative artificial intelligence, one giant has held back: Apple. The company has yet to introduce so much as an AI-generated emoji, and according to a New York Times report today and earlier reporting from Bloomberg, it is in preliminary talks with Google about adding the search company's Gemini AI model to iPhones. Yet a research paper quietly posted online last Friday by Apple engineers suggests that the company is making significant new investments into AI that are already bearing fruit. It details the development of a new generative AI model called MM1 capable of working with text and images. The researchers show it answering questions about photos and displaying the kind of general knowledge skills shown by chatbots like ChatGPT.


Microsoft hires DeepMind cofounder to lead its new consumer AI division

Engadget

Microsoft now has a lone leader overseeing consumer AI for the first time. Suleyman will try to push the consumer-facing Copilot assistant into the future, preparing for what may be a long battle with Google for artificial intelligence supremacy among Silicon Valley's Big Five companies. Suleyman's official title will be executive vice president and CEO of a new division called Microsoft AI, reporting directly to CEO Satya Nadella. Joining him will be fellow Inflection AI cofounder Karรฉn Simonyan, who takes the title of chief scientist. "Messy" could be one way to describe Microsoft's Copilot rollout.


Generative AI and CS Education

Communications of the ACM

I have spent most of my career working on computer science (CS) education whether teaching undergraduate CS or managing technical education for software engineers at Google. In the early 1990s, when Pascal was the language of choice, I began teaching CS1 and CS2 at Stanford. Over the next few years, I saw the transition from Pascal to C to object-oriented programming. I also saw the pace at which we had to consistently update our course materials and projects, whether it was in the introductory courses or later electives such as graphics or compilers. Languages, software frameworks, libraries, APIs, and so forth change rapidly.


Tech giant Nvidia unveils higher performing 'superchips' to power AI

Al Jazeera

Nvidia has unveiled its latest family of chips for powering artificial intelligence as it seeks to consolidate its position as the major supplier to the AI frenzy. So, ladies and gentlemen, I would like to introduce you to a very, very big GPU," CEO Jensen Huang said on Monday at a developers conference in California, referring to the graphics processors that are vital in the creation of generative AI. The event, dubbed the "AI Woodstock" by Wedbush analyst Dan Ives, has become a can't-miss date on big tech's calendar due to Nvidia's singular role in the AI revolution that has taken the world by storm since the introduction of ChatGPT in late 2022. "I hope you realise this is not a concert, this is a developers' conference," Huang joked as he took the stage in a packed arena usually reserved for ice hockey games and concerts. Nvidia's powerful GPU chips and software are an integral ingredient in the creation of generative AI, with rivals like AMD or Intel still struggling to match the power and efficiency of the company's blockbuster H100 product, launched in 2022.


NVIDIA's GPUs powered the AI revolution. Its new Blackwell chips are up to 30 times faster

Engadget

In less than two years, NVIDIA's H100 chips, which are used by nearly every AI company in the world to train large language models that power services like ChatGPT, made it one of the world's most valuable companies. On Monday, NVIDIA announced a next-generation platform called Blackwell, whose chips are between seven and 30 times faster than the H100 and use 25 times less power. "Blackwell GPUs are the engine to power this new Industrial Revolution," said NVIDIA CEO Jensen Huang at the company's annual GTC event in San Jose attended by thousands of developers, and which some compared to a Taylor Swift concert. "Generative AI is the defining technology of our time. Working with the most dynamic companies in the world, we will realize the promise of AI for every industry," Huang added in a press release.


AI Robots and Humanoid AI: Review, Perspectives and Directions

arXiv.org Artificial Intelligence

In the approximately century-long journey of robotics, humanoid robots made their debut around six decades ago. The rapid advancements in generative AI, large language models (LLMs), and large multimodal models (LMMs) have reignited interest in humanoids, steering them towards real-time, interactive, and multimodal designs and applications. This resurgence unveils boundless opportunities for AI robotics and novel applications, paving the way for automated, real-time and humane interactions with humanoid advisers, educators, medical professionals, caregivers, and receptionists. However, while current humanoid robots boast human-like appearances, they have yet to embody true humaneness, remaining distant from achieving human-like intelligence. In our comprehensive review, we delve into the intricate landscape of AI robotics and AI humanoid robots in particular, exploring the challenges, perspectives and directions in transitioning from human-looking to humane humanoids and fostering human-like robotics. This endeavour synergizes the advancements in LLMs, LMMs, generative AI, and human-level AI with humanoid robotics, omniverse, and decentralized AI, ushering in the era of AI humanoids and humanoid AI.


Can AI Outperform Human Experts in Creating Social Media Creatives?

arXiv.org Artificial Intelligence

Artificial Intelligence has outperformed human experts in functional tasks such as chess and baduk. How about creative tasks? This paper evaluates AI's capability in the creative domain compared to human experts, which little research has been conducted so far. We propose a novel Prompt-for-Prompt to generate social media creatives via prompt augmentation by Large Language Models. We take the most popular Instagram posts (with the biggest number of like clicks) in top brands' Instagram accounts to create social media creatives. We give GPT 4 several prompt instructions with text descriptions to generate the most effective prompts for cutting-edge text-to-image generators: Midjourney, DALL E 3, and Stable Diffusion. LLM-augmented prompts can boost AI's abilities by adding objectives, engagement strategy, lighting and brand consistency for social media image creation. We conduct an extensive human evaluation experiment, and find that AI excels human experts, and Midjourney is better than the other text-to-image generators. Surprisingly, unlike conventional wisdom in the social media industry, prompt instruction including eye-catching shows much poorer performance than those including natural. Regarding the type of creatives, AI improves creatives with animals or products but less with real people. Also, AI improves creatives with short text descriptions more than with long text descriptions, because there is more room for AI to augment prompts with shorter descriptions.


RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content

arXiv.org Artificial Intelligence

Recent advancements in Large Language Models (LLMs) have showcased remarkable capabilities across various tasks in different domains. However, the emergence of biases and the potential for generating harmful content in LLMs, particularly under malicious inputs, pose significant challenges. Current mitigation strategies, while effective, are not resilient under adversarial attacks. This paper introduces Resilient Guardrails for Large Language Models (RigorLLM), a novel framework designed to efficiently and effectively moderate harmful and unsafe inputs and outputs for LLMs. By employing a multi-faceted approach that includes energy-based training data augmentation through Langevin dynamics, optimizing a safe suffix for inputs via minimax optimization, and integrating a fusion-based model combining robust KNN with LLMs based on our data augmentation, RigorLLM offers a robust solution to harmful content moderation. Our experimental evaluations demonstrate that RigorLLM not only outperforms existing baselines like OpenAI API and Perspective API in detecting harmful content but also exhibits unparalleled resilience to jailbreaking attacks. The innovative use of constrained optimization and a fusion-based guardrail approach represents a significant step forward in developing more secure and reliable LLMs, setting a new standard for content moderation frameworks in the face of evolving digital threats.