Chapter 6. Data Analysis with Apache Pig

In the previous chapters, we explored a number of APIs for data processing. MapReduce, Spark, Tez and Samza are rather low-level, and writing non-trivial business logic with them often requires significant Java development. Moreover, different users will have different needs. It might be impractical for an analyst to write MapReduce code or build a DAG of inputs and outputs to answer some simple queries. At the same time, a software engineer or a researcher might want to prototype ideas and algorithms using high-level abstractions before jumping into low-level implementation details.

In this chapter and the following one, we will explore some tools that provide a way to process data on HDFS using higher-level ...

Get Learning Hadoop 2 now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.

Start your free trial

Learning Hadoop 2 by Garry Turkington, Gabriele Modena

Chapter 6. Data Analysis with Apache Pig

Don’t leave empty-handed

It’s yours, free.

Check it out now on O’Reilly