Guide to High Performance Distributed Computing: Case Studies with Hadoop, Scalding and Spark - eboo
This timely text/reference describes the development and implementation of large-scale distributed processing systems using open source tools and technologies such as Hadoop, Scalding and Spark.<br /><br />Comprehensive in scope, the book presents state-of-the-art material on building high performance distributed computing systems, providing practical guidance and best practices as well as describing theoretical software frameworks.<br /><br /><br /><br />Topics and features: describes the fundamentals of building scalable software systems for large-scale data processing in the new paradigm of high performance distributed computing; presents an overview of the Hadoop ecosystem, followed by step-by-step instruction on its installation, programming and execution; reviews the basics of Spark, including resilient distributed datasets, and examines Hadoop streaming and working with Scalding; provides detailed case studies on approaches to clustering, data classification and regression analy
0コメント