Goto

Collaborating Authors

 python and java


Do Code LLMs Understand Design Patterns?

arXiv.org Artificial Intelligence

Code Large Language Models (LLMs) demonstrate great versatility in adapting to various downstream tasks, including code generation and completion, as well as bug detection and fixing. However, Code LLMs often fail to capture existing coding standards, leading to the generation of code that conflicts with the required design patterns for a given project. As a result, developers must post-process to adapt the generated code to the project's design norms. In this work, we empirically investigate the biases of Code LLMs in software development. Through carefully designed experiments, we assess the models' understanding of design patterns across recognition, comprehension, and generation. Our findings reveal that biases in Code LLMs significantly affect the reliability of downstream tasks.


CAT-LM: Training Language Models on Aligned Code And Tests

arXiv.org Artificial Intelligence

Testing is an integral part of the software development process. Yet, writing tests is time-consuming and therefore often neglected. Classical test generation tools such as EvoSuite generate behavioral test suites by optimizing for coverage, but tend to produce tests that are hard to understand. Language models trained on code can generate code that is highly similar to that written by humans, but current models are trained to generate each file separately, as is standard practice in natural language processing, and thus fail to consider the code-under-test context when producing a test file. In this work, we propose the Aligned Code And Tests Language Model (CAT-LM), a GPT-style language model with 2.7 Billion parameters, trained on a corpus of Python and Java projects. We utilize a novel pretraining signal that explicitly considers the mapping between code and test files when available. We also drastically increase the maximum sequence length of inputs to 8,192 tokens, 4x more than typical code generation models, to ensure that the code context is available to the model when generating test code. We analyze its usefulness for realistic applications, showing that sampling with filtering (e.g., by compilability, coverage) allows it to efficiently produce tests that achieve coverage similar to ones written by developers while resembling their writing style. By utilizing the code context, CAT-LM generates more valid tests than even much larger language models trained with more data (CodeGen 16B and StarCoder) and substantially outperforms a recent test-specific model (TeCo) at test completion. Overall, our work highlights the importance of incorporating software-specific insights when training language models for code and paves the way to more powerful automated test generation.


Machine Learning Guide: Differences Between Python and Java

#artificialintelligence

Machine learning, Data Science, deep learning, and many other major and rising disruptive technologies need one or two programming languages to create products and services for the global tech market. Companies and start-ups have started recruiting employees who are from a technical background with sufficient knowledge of anyone or two programming languages such as Python, Java, R, C, etc. for efficient coding. Strong coding skills are essential to deal with cutting-edge technologies. Aspiring machine learning engineers, machine learning architects, data scientists, and so on are highly interested to learn Python and Java as these two have the highest demand in this field. This article is a machine learning guide for beginners to explain the differences between Python and Java for a better understanding.


10 Reasons to Ditch Java and Start Learning Python Today

#artificialintelligence

In this modern world, the technologies have shifted drastically! Some languages and technologies have seen downfall over the period of time, whereas some of them have broken the bars! Python is one of those languages that is at its peek and does not seem to stop in this race in the near future. Python is that rabbit, that is fast and furious. "Slow and steady"- is not the principle that one follows in this ever changing modern world.


Make predictions with Python machine learning for apps

#artificialintelligence

Udemy Coupon Code Link: Make predictions with Python machine learning for apps Udemy Make predictions with Python machine learning for apps. With the help of this course you can Leverage TensorFlow models to build & improve apps! What you'll learn Master the basics: become an expert in Python and Java while learning core machine learning concepts Machine learning goes mobile: learn how to incorporate machine learning models into Android apps Optimize for intelligent apps: discover the TensorFlow mobile framework and build scientific analysis apps Description Go through 3 ultimate levels of artificial intelligence for beginners! This course was funded by a wildly successful Kickstarter Use Google's deep learning framework TensorFlow with Python. Leverage machine learning to improve your apps Prediction Models Masterclass By the end of this course you will have 3 complete mobile machine learning models and apps.