SlideShare is now on Android. 15 million presentations at your fingertips.  Get the app

  • Share
  • Email
  • Embed
  • Like
  • Private Content

RHadoop, R meets Hadoop

by on Mar 06, 2012


(Presented by Antonio Piccolboni to Strata 2012 Conference, Feb 29 2012)....

(Presented by Antonio Piccolboni to Strata 2012 Conference, Feb 29 2012).

Rhadoop is an open source project spearheaded by Revolution Analytics to grant data scientists access to Hadoop’s scalability from their favorite language, R. RHadoop is comprised of three packages.

- rhdfs provides file level manipulation for HDFS, the Hadoop file system
- rhbase provides access to HBASE, the hadoop database
- rmr allows to write mapreduce programs in R

rmr allows R developers to program in the mapreduce framework, and to all developers provides an alternative way to implement mapreduce programs that strikes a delicate compromise betwen power and usability. It allows to write general mapreduce programs, offering the full power and ecosystem of an existing, established programming language. It doesn’t force you to replace the R interpreter with a special run-time—it is just a library. You can write logistic regression in half a page and even understand it. It feels and behaves almost like the usual R iteration and aggregation primitives. It is comprised of a handful of functions with a modest number of arguments and sensible defaults that combine in many useful ways. But there is no way to prove that an API works: one can only show examples of what it allows to do and we will do that covering a few from machine learning and statistics. Finally, we will discuss how to get involved.



Total Views
Views on SlideShare
Embed Views



31 Embeds 53,328 49916 2356 851 62 22 20 15 12 10 10 9 5 5 4 4 3 2 2 2 2 2 2 2 2 2 1 1 1 1 1 1


Upload Details

Uploaded via SlideShare as Apple Keynote

Usage Rights

© All Rights Reserved

Report content

Flagged as inappropriate Flag as inappropriate
Flag as inappropriate

Select your reason for flagging this presentation as inappropriate.


12 of 2 previous next

  • halueda Haruyasu Ueda, 主任研究員/Senior researcher at Fujitsu Hmmm. I'm not sure the difference between RHadoop and RHive. RHive has mapreduce functionality even though its name is Hive. It also has HDFS adapter. 1 year ago
    Are you sure you want to
    Your message goes here
  • sildershare2010 学峰 司 RHadoop step by step 1 year ago
    Are you sure you want to
    Your message goes here
Post Comment
Edit your comment

RHadoop, R meets Hadoop RHadoop, R meets Hadoop Presentation Transcript