sphinx
Boosting Skeleton-Driven SMT Solver Fuzzing by Leveraging LLM to Produce Formula Generators
Sun, Maolin, Yang, Yibiao, Zhou, Yuming
Satisfiability Modulo Theory (SMT) solvers are foundational to modern systems and programming languages research, providing the foundation for tasks like symbolic execution and automated verification. Because these solvers sit on the critical path, their correctness is essential, and high-quality test formulas are key to uncovering bugs. However, while prior testing techniques performed well on earlier solver versions, they struggle to keep pace with rapidly evolving features. Recent approaches based on Large Language Models (LLMs) show promise in exploring advanced solver capabilities, but two obstacles remain: nearly half of the generated formulas are syntactically invalid, and iterative interactions with the LLMs introduce substantial computational overhead. In this study, we present Chimera, a novel LLM-assisted fuzzing framework that addresses both issues by shifting from direct formula generation to the synthesis of reusable term (i.e., logical expression) generators. Particularly, Chimera uses LLMs to (1) automatically extract context-free grammars (CFGs) for SMT theories, including solver-specific extensions, from documentation, and (2) synthesize composable Boolean term generators that adhere to these grammars. During fuzzing, Chimera populates structural skeletons derived from existing formulas with the terms iteratively produced by the LLM-synthesized generators. This design ensures syntactic validity while promoting semantic diversity. Notably, Chimera requires only one-time LLM interaction investment, dramatically reducing runtime cost. We evaluated Chimera on two leading SMT solvers: Z3 and cvc5. Our experiments show that Chimera has identified 43 confirmed bugs, 40 of which have already been fixed by developers.
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
Ahuja, Sanchit, Tanmay, Kumar, Chauhan, Hardik Hansrajbhai, Patra, Barun, Aggarwal, Kriti, Del Corro, Luciano, Mitra, Arindam, Dhamecha, Tejas Indulal, Awadallah, Ahmed, Choudhary, Monojit, Chaudhary, Vishrav, Sitaram, Sunayana
Despite the remarkable success of LLMs in English, there is a significant gap in performance in non-English languages. In order to address this, we introduce a novel recipe for creating a multilingual synthetic instruction tuning dataset, sPhinX, which is created by selectively translating instruction response pairs from English into 50 languages. We test the effectiveness of sPhinX by using it to fine-tune two state-of-the-art models, Phi-3-small and Mistral-7B and then evaluating them across a comprehensive suite of multilingual benchmarks that test reasoning, question answering, and reading comprehension. Our results show that Phi-3-small and Mistral-7B fine-tuned with sPhinX perform better on an average by 4.2%pt and 5%pt respectively as compared to the baselines. We also devise a strategy to incorporate N-shot examples in each fine-tuning sample which further boosts the performance of these models by 3%pt and 10%pt respectively. Additionally, sPhinX also outperforms other multilingual instruction tuning datasets on the same benchmarks along with being sample efficient and diverse, thereby reducing dataset creation costs. Additionally, instruction tuning with sPhinX does not lead to regression on most standard LLM benchmarks.
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Lin, Ziyi, Liu, Chris, Zhang, Renrui, Gao, Peng, Qiu, Longtian, Xiao, Han, Qiu, Han, Lin, Chen, Shao, Wenqi, Chen, Keqin, Han, Jiaming, Huang, Siyuan, Zhang, Yichi, He, Xuming, Li, Hongsheng, Qiao, Yu
We present SPHINX, a versatile multi-modal large language model (MLLM) with a joint mixing of model weights, tuning tasks, and visual embeddings. First, for stronger vision-language alignment, we unfreeze the large language model (LLM) during pre-training, and introduce a weight mix strategy between LLMs trained by real-world and synthetic data. By directly integrating the weights from two domains, the mixed LLM can efficiently incorporate diverse semantics with favorable robustness. Then, to enable multi-purpose capabilities, we mix a variety of tasks for joint visual instruction tuning, and design task-specific instructions to avoid inter-task conflict. In addition to the basic visual question answering, we include more challenging tasks such as region-level understanding, caption grounding, document layout detection, and human pose estimation, contributing to mutual enhancement over different scenarios. Additionally, we propose to extract comprehensive visual embeddings from various network architectures, pre-training paradigms, and information granularity, providing language models with more robust image representations. Based on our proposed joint mixing, SPHINX exhibits superior multi-modal understanding capabilities on a wide range of applications. On top of this, we further propose an efficient strategy aiming to better capture fine-grained appearances of high-resolution images. With a mixing of different scales and high-resolution sub-images, SPHINX attains exceptional visual parsing and reasoning performance on existing evaluation benchmarks. We hope our work may cast a light on the exploration of joint mixing in future MLLM research. Code is released at https://github.com/Alpha-VLLM/LLaMA2-Accessory.
7 Python Tools Every ML Developer & Data Scientist Should Have
Python is a popular programming language that has become the favored option for software developers and data scientists alike, from constructing advanced machine learning algorithms to creating easy graphical user interfaces. Python's data science skills are still being explored, especially for advanced data analysis and the creation of deep learning solutions. In this approach, Python beats other programming languages such as C . Python has a modest learning curve and is considered very beginner-friendly. But many tools must be understood to obtain the maximum benefit from Python.
3 Python Tools Data Scientists Can Use for Production-Quality Code
For many of these steps, there are no real short cuts to be taken. The only way to build a minimum viable product, for example, is to roll up your sleeves and start coding. However, in a few cases, tools exist to automate tedious manual processes and make your life much easier. In Python, this is the situation for steps 4, 8 and 10, thanks to the unittest, flake8 and sphinx packages. Let's look at each of these packages one by one.
How Smart Is Artificial Intelligence?
Artificial intelligence (AI) has come a long way in recent years. For example, AI has defeated the human world champion of Go, recreated the periodic table of elements, enabled self-driving vehicles, identifed crop diseases, and predicted depression from speech. Imagine what could happen as AI improves capabilities in areas that are squarely in the domain of the human brain. Today, technology is far from achieving parity with human-level intelligence, also known as "strong AI" or artificial general intelligence (AGI). Recent AI advances in one capability--pattern recognition--has spawned an investment gold rush for AI startups and machine learning talent from venture capital, corporations, and governments who have recognized the potential competitive advantage.
'Westworld' Is Turning Into Lost--for Better or for Worse
I never should have started watching Westworld. Not because I didn't think it'd be good. An HBO show based on a Michael Crichton idea starring Evan Rachel Wood with all kinds of artificial intelligence? The problem wasn't that Westworld wouldn't be enjoyable, it was that it's the kind of show that invites obsession. The kind that presents Big Questions--that never get answered.
'Assassin's Creed Origins' virtual tours can actually teach history
The Assassin's Creed series is known for its vast and richly detailed historical environments, and well... lots of murder. What you might not realize is just how much work goes into making these virtual windows into the past somewhat realistic. That's something Ubisoft is aiming to highlight with Assassin's Creed Origins' Discovery Tour. You can think of it as a museum-like experience set within the game's meticulous rendition of ancient Egypt. To turn one of the most popular gaming franchises in the world into a truly useful educational tool.
How Assassin's Creed: Origins' photo mode makes your Egyptian tourism feel important
For years I've claimed Assassin's Creed is better as virtual tourism than it is as a game proper. It was only a matter of time until they gave players a camera. Uplay tells me I'm at 83 percent completion overall--just some forts and a few side quests to mop up, but I've finished Bayek's story, learned how the Assassins came into existence, uncovered the secrets of the Pyramids and the Great Sphinx. And Bayek's right, the Sphinx is smaller than I thought. And yet my map is still full of icons.
Coding Jarvis in Python in 2016
It's tough for an erstwhile Iron Man to work on creating their personal AI assistant on the weekends. Like any other time-pressured inventor without a PhD in computer science and linguistics, I decided to use a library for speech recognition and synthesis. Fortunately, Python offers several choices. I will discuss the ones that are still functional and can be used with Python 2.7 and Python 3 (up to Python 3.5 at the time of writing). We've all heard of Google's AI initiatives (maybe Larry Page aspires to Tony Stark status?), so it should come as little surprise that they offer a RESTful way to do voice recognition and speech synthesis.