Hadoop Operational

NameNode and DataNodes

https://hadoop.apache.org/docs/r1.0.4/hdfs_design.html HDFS has a master/slave architecture. An HDFS cluster consists of a single NameNode, a master server that manages the file system namespace and regulates access to files by clients. In addition, there are a number of DataNodes, usually one per node in the cluster, which manage storage attached to the nodes that they run on. HDFS exposes a file system namespace and allows user data to be stored in files. Internally, a file is split into one or more blocks and these blocks are stored in a set of DataNodes. The NameNode executes file system namespace operations like opening, closing, and renaming files and directories. It also determines the mapping of blocks to DataNodes. The DataNodes are responsible for serving read and write requests from the file system’s clients. The DataNodes also perform block creation, deletion, and replication upon instruction from the NameNode.

File System

The File System Namespace HDFS supports a traditional hierarchical file organization. A user or an application can create directories and store files inside these directories. The file system namespace hierarchy is similar to most other existing file systems; one can create and remove files, move a file from one directory to another, or rename a file. HDFS does not yet implement user quotas. HDFS does not support hard links or soft links. However, the HDFS architecture does not preclude implementing these features. The NameNode maintains the file system namespace. Any change to the file system namespace or its properties is recorded by the NameNode. An application can specify the number of replicas of a file that should be maintained by HDFS. The number of copies of a file is called the replication factor of that file. This information is stored by the NameNode.

Configuration

Data ingest and replication

How data is saved to HDFS

Lab: HDFS ingest and replication

Copy a file into HDFS with 1 MB (1048576 b) blocksize
Copy a file into HDFS with a 5x replication factor
Look up all the information that’s available on the webpages for the services

NOTE: Complete this work on the Gateway Node

Lab Solution


hadoop fs -D dfs.blocksize=30 -put somefile somelocation
hdfs fsck filelocation -files -blocks -locations
hdfs dfs -mkdir test
hdfs dfs -ls

Namenode functionality

Keeps track of the HDFS namespace metadata
Controls opening, closing, and renaming files/directories by clients

Datanode functionality

Handles read/write requests for blocks
Reports blocks back to the Namenode
Replicates blocks to other datanodes

File Permission

Supports POSIX-style permissions

user:group ownership
rwx permissions on user, group, and other


$ hdfs dfs -ls myfile
-rw-r--r--   1 swanftw swanftw          1 2014-07-24 20:14 myfile
$ hdfs dfs -chmod 755 myfile
$ hdfs dfs -ls myfile
-rwxr-xr-x   1 swanftw swanftw      	1 2014-07-24 20:14 myfile

File System Shell

File system shell (HDFS client)

The hdfs dfs -rm command allows you to remove files from HDFS. It supports additional arguments such as -r for recursive deletes and -skipTrash if you’d like to permanently delete files rather than moving them to the trash. hdfs dfs -put can be used to upload a file from the local system. The following two arguments are the local file followed by the HDFS filename and/or path. By using a dash in place of the local filename, you can read data from stdin instead of from a file. This allows you to pipe data into HDFS, such as output from echo as in the example. Ownership and permissions are modified using mdfs dfs -chown and -chmod, with usage almost identical to the chown and chmod commands on *nix systems. hdfs dfs -cat allow you to output the contents of files in HDFS. This contents can be displayed in the terminal or piped into other commands as needed.

File system shell (HDFS client) cont...