Government
A robust synthetic data generation framework for machine learning in High-Resolution Transmission Electron Microscopy (HRTEM)
DaCosta, Luis Rangel, Sytwu, Katherine, Groschner, Catherine, Scott, Mary
Machine learning techniques are attractive options for developing highly-accurate automated analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapidly generating complex nanoscale atomic structures, and develop an end-to-end workflow for creating large simulated databases for training neural networks. Construction Zone enables fast, systematic sampling of realistic nanomaterial structures, and can be used as a random structure generator for simulated databases, which is important for generating large, diverse synthetic datasets. Using HRTEM imaging as an example, we train a series of neural networks on various subsets of our simulated databases to segment nanoparticles and holistically study the data curation process to understand how various aspects of the curated simulated data -- including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions -- affect model performance across several experimental benchmarks. Using our results, we are able to achieve state-of-the-art segmentation performance on experimental HRTEM images of nanoparticles from several experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data.
Digital Twin System for Home Service Robot Based on Motion Simulation
Jiang, Zhengsong, Tian, Guohui, Cui, Yongcheng, Liu, Tiantian, Gu, Yu, Wang, Yifei
In order to improve the task execution capability of home service robot, and to cope with the problem that purely physical robot platforms cannot sense the environment and make decisions online, a method for building digital twin system for home service robot based on motion simulation is proposed. A reliable mapping of the home service robot and its working environment from physical space to digital space is achieved in three dimensions: geometric, physical and functional. In this system, a digital space-oriented URDF file parser is designed and implemented for the automatic construction of the robot geometric model. Next, the physical model is constructed from the kinematic equations of the robot and an improved particle swarm optimization algorithm is proposed for the inverse kinematic solution. In addition, to adapt to the home environment, functional attributes are used to describe household objects, thus improving the semantic description of the digital space for the real home environment. Finally, through geometric model consistency verification, physical model validity verification and virtual-reality consistency verification, it shows that the digital twin system designed in this paper can construct the robot geometric model accurately and completely, complete the operation of household objects successfully, and the digital twin system is effective and practical.
Learning Unbiased News Article Representations: A Knowledge-Infused Approach
Kamal, Sadia, Hartford, Jimmy, Willis, Jeremy, Bagavathi, Arunkumar
Quantification of the political leaning of online news articles can aid in understanding the dynamics of political ideology in social groups and measures to mitigating them. However, predicting the accurate political leaning of a news article with machine learning models is a challenging task. This is due to (i) the political ideology of a news article is defined by several factors, and (ii) the innate nature of existing learning models to be biased with the political bias of the news publisher during the model training. There is only a limited number of methods to study the political leaning of news articles which also do not consider the algorithmic political bias which lowers the generalization of machine learning models to predict the political leaning of news articles published by any new news publishers. In this work, we propose a knowledge-infused deep learning model that utilizes relatively reliable external data resources to learn unbiased representations of news articles using their global and local contexts. We evaluate the proposed model by setting the data in such a way that news domains or news publishers in the test set are completely unseen during the training phase. With this setup we show that the proposed model mitigates algorithmic political bias and outperforms baseline methods to predict the political leaning of news articles with up to 73% accuracy.
ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
Golovneva, Olga, Chen, Moya, Poff, Spencer, Corredor, Martin, Zettlemoyer, Luke, Fazel-Zarandi, Maryam, Celikyilmaz, Asli
Large language models show improved downstream task performance when prompted to generate step-by-step reasoning to justify their final answers. These reasoning steps greatly improve model interpretability and verification, but objectively studying their correctness (independent of the final answer) is difficult without reliable methods for automatic evaluation. We simply do not know how often the stated reasoning steps actually support the final end task predictions. In this work, we present ROSCOE, a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics. To evaluate ROSCOE against baseline metrics, we design a typology of reasoning errors and collect synthetic and human evaluation scores on commonly used reasoning datasets. In contrast with existing metrics, ROSCOE can measure semantic consistency, logicality, informativeness, fluency, and factuality - among other traits - by leveraging properties of step-by-step rationales. We empirically verify the strength of our metrics on five human annotated and six programmatically perturbed diagnostics datasets - covering a diverse set of tasks that require reasoning skills and show that ROSCOE can consistently outperform baseline metrics.
The Real Stakes of the Google Antitrust Trial
The year 1998 was a pivotal one in the history of technology: Apple's introduction of the iMac helped set the company back on the path to success after it nearly went bankrupt earlier in the decade; Google was founded by two Stanford students, Larry Page and Sergey Brin; and Microsoft introduced Windows 98, an improved version of its popular computer operating system. That May, Microsoft also became the target of a historic antitrust lawsuit lodged by the Department of Justice and twenty states, accusing it of anticompetitive behavior in two domains: attempting to maintain its monopoly in computer operating systems and trying to monopolize a new market, that of Internet browsers. At the time, residential Wi-Fi connectivity was rapidly expanding across America, and, in the quaintly titled "browser wars," Netscape Navigator, a popular browser released by Mosaic Communications Corporation in 1994, fought Microsoft's Internet Explorer for the growing class of Web-connected consumers. Microsoft, the D.O.J. alleged, had attempted to crush Netscape by making deals with Internet-service providers that prioritized Explorer access at Netscape users' expense. The trial began that fall, and included seventy-six days of testimony that took place over more than eight months, during which a government witness alleged that a Microsoft executive had pledged to "cut off Netscape's air supply" (which a Microsoft attorney denied).
Kamala Harris taken aback by CBS host asking about Trump's re-election hopes: 'Don't understand the question'
Vice President Harris appeared stunned by CBS' Margaret Brennan's question about whether she was taking the threat of another Trump presidency "seriously enough." Vice President Kamala Harris appeared stunned by a question from CBS host Margaret Brennan, who wondered if she was taking the possibility of another Donald Trump presidency "seriously enough." Harris looked taken aback, pausing before responding, "I don't understand the question." "You were dismissive of some of the Republican criticism of you and the president. When you look at current polling, the frontrunner for the Republican nomination is the former president, the 45th president," Brennan added on "Face The Nation."
The Download: what to expect from US Congress's first AI meeting
The US Congress is heading back into session, and they're hitting the ground running on AI. We're going to be hearing a lot about various plans and positions on AI regulation in the coming weeks, kicking off with Senate Majority Leader Chuck Schumer's first AI Insight Forum on Wednesday. This and planned future forums will bring together some of the top people in AI to discuss the risks and opportunities it poses and how Congress might write legislation to address them. Although the forums are closed to the public and press, our senior tech policy reporter Tate Ryan-Mosley has chatted with representatives from attendee AI company Hugging Face about what they are expecting, and what exactly these forums are hoping to achieve. Tate's story first appeared in The Technocrat, her weekly newsletter covering policy and Silicon Valley.
Will Anyone Ever Make Sense of Elon Musk?
Elon Musk is "wired for war." At least, that's what Musk has told Walter Isaacson, whose thick biography of the mercurial mega-billionaire, Elon Musk, is out this week. When Musk says this, he's not talking about Ukraine, where his Starlink internet service has played a central role. Civilization, Warcraft: Orcs & Humans, The Battle of Polytopia, Elden Ring--Musk has spent much of his life in fantasy worlds. Isaacson's biography includes many astonishing details and relatively few pages focused on Musk's gaming obsession. But the video-game detail is telling. Musk doesn't seem to inhabit our reality, exactly, even as he profoundly shapes it.
AI Chatbots Are Invading Your Local Government--and Making Everyone Nervous
The United States Environmental Protection Agency blocked its employees from accessing ChatGPT while the US State Department staff in Guinea used it to draft speeches and social media posts. Maine banned its executive branch employees from using generative artificial intelligence for the rest of the year out of concern for the state's cybersecurity. In nearby Vermont, government workers are using it to learn new programming languages and write internal-facing code, according to Josiah Raiche, the state's director of artificial intelligence. The city of San Jose, California, wrote 23 pages of guidelines on generative AI and requires municipal employees to fill out a form every time they use a tool like ChatGPT, Bard, or Midjourney. Less than an hour's drive north, Alameda County's government has held sessions to educate employees about generative AI's risks--such as its propensity for spitting out convincing but inaccurate information--but doesn't see the need yet for a formal policy.
What to know about Congress's inaugural AI meeting
This newsletter will break down what exactly these forums are and aren't, and what might come out of them. The forums will be closed to the public and press, so I chatted with people at one company--Hugging Face--that did get the invite about what they are expecting and what their priorities are heading into the discussions. Schumer first announced the forums at the end of June as part of his AI legislation initiative, called SAFE Innovation. In floor remarks on Tuesday, Schumer said he's planning for "an open discussion about how Congress can act on AI: where to start, what questions to ask, and how to build a foundation for SAFE AI innovation." The SAFE framework, as a reminder, is not a legislative proposal but rather a set of priorities that Schumer laid out when it comes to AI regulation.