Ricardo: integrating R and Hadoop
Explore this paper's citation graph
Summary
R Ricardo is part of the eXtreme Analytics Platform (XAP) project at the IBM Almaden Research Center, and rests on a decomposition of data-analysis algorithms into parts executed by the R statistical analysis system and parts handled by the Hadoop data management system.
- Type
- article
- Published
- 2010-06-06
- Cited by
- 211
- References
- 25
- OpenAlex
- https://openalex.org/W2013373704
- Semantic Scholar
- https://api.semanticscholar.org/CorpusID:2470829
Keywords
Petabyte, Computer science, IBM, Terabyte, Scalability
References
- Requirements for Science Data Bases and SciDB
- Introduction to recommender systems
- Modeling relationships at multiple scales to improve accuracy of large recommender systems
- Factorization meets the neighborhood: a multifaceted collaborative filtering model
- A Limited Memory Algorithm for Bound Constrained Optimization
- Matrix Factorization Techniques for Recommender Systems
- Collaborative filtering with privacy via factor analysis
- Collaborative filtering with temporal dynamics
- Pig latin: a not-so-foreign language for data processing
- Hive - A Warehousing Solution Over a Map-Reduce Framework
- Up close and personalized: a marketing view of recommendation systems
- MAD Skills: New Analysis Practices for Big Data
- State of the Art in Parallel Computing with R
- MapReduce: simplified data processing on large clusters
- The Netflix Prize
- Nonlinear Programming: Analysis and Methods
- RIOT: I/O-Efficient Numerical Computing without SQL
- Map-Reduce for Machine Learning on Multicore
Cited by
- Performance Modeling And Resource Management For Mapreduce Applications
- Optimizing Data Movement in Hybrid Analytic Systems
- Association affinity network based multi-model collaboration for multimedia big data management and retrieval
- Big Data with Cloud Computing: an insight on the computing environment, MapReduce, and programming frameworks
- A Parallel Distributed Weka Framework for Big Data Mining Using Spark
- CBA: Cloud-Based Bigdata Analytics
- Hadoop as Big Data Operating System -- The Emerging Approach for Managing Challenges of Enterprise Big Data Platform
- Jaql
- Research on performance optimization and visualization tool of Hadoop
- Cloud Computing and Big Data Analytics
- A first view of exedra: a domain-specific language for large graph analytics workflows
- Parameterizable benchmarking framework for designing a MapReduce performance model
- A Scalable Data Science Workflow Approach for Big Data Bayesian Network Learning
- Towards Integrated Data Analytics: Time Series Forecasting in DBMS
- Presto: distributed machine learning and graph processing with sparse matrices
- LINVIEW: incremental view maintenance for complex analytical queries
- An Optimized Method for Access of LOBs in Database Management Systems
- Large-scale matrix factorization with distributed stochastic gradient descent
- RABID: A Distributed Parallel R for Large Datasets
- PARALLEL IMAGE DATABASE PROCESSING WITH MAPREDUCE AND PERFORMANCE EVALUATION IN PSEUDO DISTRIBUTED MODE
Related papers
- Lessons Learned from Managing a Petabyte
- Introduction to Big Data: Scalable Representation and Analytics for Data Science Minitrack
- BIG DATA - IMPORTANCE OF HADOOP DISTRIBUTED FILESYSTEM
- Research Data Publication at Large Scale
- Managing Large Scale Data for Earthquake Simulations
- Large scale research data archiving: Training for an inconvenient technology
- Plenary talk: Big data and real time analytics