Skip to content
KnowledgeCity

Big Data Systems and Analytics

Learn about Big Data processing systems.
Preview the first lesson free — get full access to all 4 lessons.
Course: On-Demand
Beginner Provider Chintan Thakkar  4 Lessons ·  18m  in Arabic, German, English, Spanish, French, Portuguese, Urdu, Chinese 

Course Description

To process Big Data, you need a centralized application that allows you to submit then store data into a relational database. In this Big Data Systems and Analytics course you will learn about storage systems designed for Big Data like Hadoop and its architecture, which includes Hadoop Distributed File System, MapReduce, and different types of available nodes.

The Hadoop Distributed File System provides reliability, scalability, and availability. With MapReduce, Hadoop executes the job using the data stored on the Distributed File System and helps users make computations. Apache Pig is built on top of MapReduce algorithms. It helps users create the data analysis program and analyze large datasets. This course will explain the process of how to run Big Data anyltics on Hadoop.

What You'll Learn

  • Identify the functions of MapReduce algorithms
  • Explain what the Hadoop Distributed File System (HDFS) is and its relation to Big Data
  • List the most common uses of Apache Pig
  • Recall the nodes in the HDFS architecture
  • Explain the steps for creating files in HDFS
  • Describe how to run Big Data analytics on Hadoop

Key Takeaways

  • Processing Big Data requires a centralized application that lets you submit and store data into a relational database.
  • Hadoop is a storage system designed for Big Data, with an architecture that includes the Hadoop Distributed File System, MapReduce, and different types of available nodes.
  • The Hadoop Distributed File System provides reliability, scalability, and availability.
  • With MapReduce, Hadoop executes jobs using data stored on the Distributed File System and helps users make computations.
  • Apache Pig is built on top of MapReduce algorithms and helps users create data analysis programs and analyze large datasets.

Frequently Asked Questions

What does this course cover?

The course covers storage systems designed for Big Data like Hadoop and its architecture, including the Hadoop Distributed File System, MapReduce, and different types of available nodes, as well as Apache Pig and the process of running Big Data analytics on Hadoop.

What is the Hadoop Distributed File System and why does it matter for Big Data?

The Hadoop Distributed File System provides reliability, scalability, and availability, and the course explains what it is and its relation to Big Data.

What is Apache Pig used for according to this course?

Apache Pig is built on top of MapReduce algorithms and helps users create data analysis programs and analyze large datasets; the course also lists its most common uses.

What skills will I gain from this course?

You will gain skills in Big Data, Big Data Analytics, Data Processing Systems, Database Systems, the Hadoop Distributed File System (HDFS), and Oracle Big Data.

What topics are taught in the lessons?

The lessons include Introduction to Hadoop, Hadoop Architecture, How to Run Big Data Analytics on Hadoop and HDFS, and the MapReduce Algorithm.

Professional Certifications and Continuing Education Units (CEUs)

Society for Human Resource Management (SHRM®)

Professional Development Credits (PDCs): 0.5

Certification Program Categories:
Leadership & NavigationBusiness AcumenConsultationGlobal MindsetEthical PracticeRelationship ManagementAnalytical AptitudeCommunicationDiversity, Equity & Inclusion

KnowledgeCity is approved by SHRM as a Recertification General Provider to offer SHRM-CP or SHRM-SCP professional development credits (PDCs). By taking the courses approved by SHRM, KnowledgeCity can award SHRM Professional Development Credits (PDCs) for HR knowledge and competency programs related to the SHRM Body of Applied Skills and Knowledge™ (the SHRM BASK™).