• Email
  • Like
  • Save
  • Private Content
  • Embed
 

RHadoop, R meets Hadoop

by on Mar 06, 2012

  • 38,308 views

(Presented by Antonio Piccolboni to Strata 2012 Conference, Feb 29 2012)....

(Presented by Antonio Piccolboni to Strata 2012 Conference, Feb 29 2012).

Rhadoop is an open source project spearheaded by Revolution Analytics to grant data scientists access to Hadoop’s scalability from their favorite language, R. RHadoop is comprised of three packages.

- rhdfs provides file level manipulation for HDFS, the Hadoop file system
- rhbase provides access to HBASE, the hadoop database
- rmr allows to write mapreduce programs in R

rmr allows R developers to program in the mapreduce framework, and to all developers provides an alternative way to implement mapreduce programs that strikes a delicate compromise betwen power and usability. It allows to write general mapreduce programs, offering the full power and ecosystem of an existing, established programming language. It doesn’t force you to replace the R interpreter with a special run-time—it is just a library. You can write logistic regression in half a page and even understand it. It feels and behaves almost like the usual R iteration and aggregation primitives. It is comprised of a handful of functions with a modest number of arguments and sensible defaults that combine in many useful ways. But there is no way to prove that an API works: one can only show examples of what it allows to do and we will do that covering a few from machine learning and statistics. Finally, we will discuss how to get involved.

Accessibility

Categories

Upload Details

Uploaded via SlideShare as Apple Keynote

Usage Rights

© All Rights Reserved

Flagged as inappropriate Flag as inappropriate
Flag as inappropriate

Select your reason for flagging this presentation as inappropriate. If needed, use the feedback form to let us know more details.

Cancel

20 Embeds 34,343

http://blog.revolutionanalytics.com 31264
http://www.r-bloggers.com 2267
http://smartdatacollective.com 660
http://www.scoop.it 43
http://feeds.feedburner.com 20
http://translate.googleusercontent.com 19
http://www.newsblur.com 15
http://atomicules.co.uk 12
http://webcache.googleusercontent.com 9
http://vizdat.collected.info 9
http://core.traackr.com 5
http://03.collected.info 4
http://127.0.0.1 3
http://staffmail.scu.edu.au 2
http://www.twylah.com 2
http://revolution-computing.typepad.com 2
http://www.hanrss.com 2
http://xianguo.com 2
http://webmail.scu.edu.au 2
http://cache.baidu.com 1

More...

Statistics

Likes
20
Downloads
406
Comments
2
Embed Views
34,343
Views on SlideShare
3,965
Total Views
38,308

12 of 2 previous next

Post Comment
Edit your comment

RHadoop, R meets Hadoop RHadoop, R meets Hadoop Presentation Transcript