alluxio
Orchestrating data for machine learning pipelines
Machine learning (ML) workloads require efficient infrastructure to yield rapid results. Model training relies heavily on large data sets. Funneling this data from storage to the training cluster is the first step of any ML workflow, which significantly impacts the efficiency of model training. This article will discuss a new solution to orchestrating data for end-to-end machine learning pipelines that addresses the above questions. I will outline common challenges and pitfalls, followed by proposing a new technique, data orchestration, to optimize the data pipeline for machine learning.
Machine Learning Magic: How to Speed Up Offline Inference for Large Datasets
In this blog, guest writers Binyang Li (Software Engineer at Microsoft), Qianxi Zhang (Research Software Engineer at Microsoft), describe how to use Alluxio to solve the challenges while running inference at scale. The original content was published on Alluxio's Blog (Disclaimer: The author is a Founding Member @Alluxio). Offline inference, or batch inference, is an approach to run machine learning (ML) inference in a batch mode when processing a large dataset, as opposed to generating predictions in real-time given the input. The offline inference jobs are typically built on top of big data platforms to scale horizontally, and are running on fixed schedules (e.g. Running inference at scale is challenging.
Alluxio Raises $50 Million In Funding, Launches New Release Of Its Data Orchestration Platform
Alluxio has raised $50 million in a Series C round of funding, capital the company will use to fuel the growth of its global operations and continue building out the capabilities of its data orchestration software for managing large-scale distributed data workloads. With the additional capital Alluxio will "enlarge our bandwidth in research and development to expand product capabilities, as well as increase our go-to-market capacity in different regions," said founder and CEO Haoyuan Li in an interview with CRN. The company is particularly looking to expand its Asia-Pacific presence and just opened an office in Beijing, China. Alluxio also announced the availability of version 2.7 of its Data Orchestration Platform with improved I/O performance for machine learning and support for open table formats such as Apachi Hudi and Iceberg. Alluxio's software, a virtual distributed file system that separates compute from storage, provides a way to unify access to data scattered across widely distributed hybrid-cloud and multi-cloud environments, making all data appear local no matter where it's stored.
How big data and AI work together
Big data isn't quite the term de rigueur that it was a few years ago, but that doesn't mean it went anywhere. If anything, big data has just been getting bigger. That once might have been considered a significant challenge. But now, it's increasingly viewed as a desired state, specifically in organizations that are experimenting with and implementing machine learning and other AI disciplines. "AI and ML are now giving us new opportunities to use the big data that we already had, as well as unleash a whole lot of new use cases with new data types," says Glenn Gruber, senior digital strategist at Anexinet. "We now have much more usable data in the form of pictures, video, and voice [for example].
How big data and AI work together
Big data isn't quite the term de rigueur that it was a few years ago, but that doesn't mean it went anywhere. If anything, big data has just been getting bigger. That once might have been considered a significant challenge. But now, it's increasingly viewed as a desired state, specifically in organizations that are experimenting with and implementing machine learning and other AI disciplines. "AI and ML are now giving us new opportunities to use the big data that we already had, as well as unleash a whole lot of new use cases with new data types," says Glenn Gruber, senior digital strategist at Anexinet. "We now have much more usable data in the form of pictures, video, and voice [for example].
In the age of AI, fundamental value resides in data
The call for speakers for our 2019 Artificial Intelligence Conference in Beijing, China, closes at midnight (China time) on January 8, 2019. Subscribe to the O'Reilly Data Show Podcast to explore the opportunities and techniques driving big data, data science, and AI. Find us on Stitcher, TuneIn, iTunes, SoundCloud, RSS. In this episode of the Data Show, I spoke with Haoyuan Li, CEO and founder of Alluxio, a startup commercializing the open source project with the same name (full disclosure: I'm an advisor to Alluxio). Our discussion focuses on the state of Alluxio (the open source project that has roots in UC Berkeley's AMPLab), specifically emerging use cases here and in China.
Why businesses should pay attention to deep learning
In this episode of the O'Reilly Data Show, I spoke with Christopher Nguyen, CEO and co-founder of Arimo. Nguyen and Arimo were among the first adopters and proponents of Apache Spark, Alluxio, and other open source technologies. Most recently, Arimo's suite of analytic products has relied on deep learning to address a range of business problems. When we started Arimo (our company name then was Adatao), the vision was about big data and machine learning. At the time, the industry had just refactored itself into what I call the'big data layer'--big data in the sense of the layer at the bottom, the storage layer.