Cover by Chuck Lam

Safari, the world’s most comprehensive technology and business learning platform.

Find the exact information you need to solve a problem on the fly, or go deeper to master the technologies and skills you need to succeed

Start Free Trial

No credit card required

O'Reilly logo

Chapter 11. Hive and the Hadoop herd

This chapter covers

  • What Hive is
  • Setting up Hive
  • Using Hive for data warehousing
  • Other software packages related to Hadoop

As powerful as Hadoop is, it doesn’t offer everything for everybody. Many projects have sprung up to extend Hadoop for specific purposes. The most prominent and well-supported ones have officially become subprojects under the umbrella of the Apache Hadoop project.[1] These subprojects include

1 What we’ve referred to in this book as “Hadoop” so far (HDFS and MapReduce) is technically called the “Hadoop Core” subproject of Apache Hadoop, although colloquially people tend to call it Hadoop.

  • Pig— A high-level data flow language
  • Hive— A SQL-like data warehouse infrastructure

Find the exact information you need to solve a problem on the fly, or go deeper to master the technologies and skills you need to succeed

Start Free Trial

No credit card required