Apache Pig-データの保存

前の章で、Apache Pigにデータをロードする方法を学びました。 store 演算子を使用して、ロードされたデータをファイルシステムに保存できます。この章では、 Store 演算子を使用してApache Pigにデータを保存する方法について説明します。

構文

以下に、Storeステートメントの構文を示します。

STORE Relation_name INTO ' required_directory_path ' [USING function];

例

次の内容のファイル student_data.txt がHDFSにあると仮定します。

001,Rajiv,Reddy,9848022337,Hyderabad
002,siddarth,Battacharya,9848022338,Kolkata
003,Rajesh,Khanna,9848022339,Delhi
004,Preethi,Agarwal,9848022330,Pune
005,Trupthi,Mohanthy,9848022336,Bhuwaneshwar
006,Archana,Mishra,9848022335,Chennai.

そして、以下に示すように、LOAD演算子を使用して関係*学生*に読み込みました。

grunt> student = LOAD 'hdfs://localhost:9000/pig_data/student_data.txt'
   USING PigStorage(',')
   as ( id:int, firstname:chararray, lastname:chararray, phone:chararray,
   city:chararray );

次に、以下に示すように、HDFSディレクトリ*“/pig_Output/” *にリレーションを保存します。

grunt> STORE student INTO ' hdfs://localhost:9000/pig_Output/' USING PigStorage (',');

出力

*store* ステートメントを実行すると、次の出力が得られます。 指定した名前でディレクトリが作成され、データがそこに保存されます。

2015-10-05 13:05:05,429 [main] INFO  org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.
MapReduceLau ncher - 100% complete
2015-10-05 13:05:05,429 [main] INFO  org.apache.pig.tools.pigstats.mapreduce.SimplePigStats -
Script Statistics:

HadoopVersion    PigVersion    UserId    StartedAt             FinishedAt             Features
2.6.0            0.15.0        Hadoop    2015-10-0 13:03:03    2015-10-05 13:05:05    UNKNOWN
Success!
Job Stats (time in seconds):
JobId          Maps    Reduces    MaxMapTime    MinMapTime    AvgMapTime    MedianMapTime
job_14459_06    1        0           n/a           n/a           n/a           n/a
MaxReduceTime    MinReduceTime    AvgReduceTime    MedianReducetime    Alias    Feature
     0                 0                0                0             student  MAP_ONLY
OutPut folder
hdfs://localhost:9000/pig_Output/

Input(s): Successfully read 0 records from: "hdfs://localhost:9000/pig_data/student_data.txt"
Output(s): Successfully stored 0 records in: "hdfs://localhost:9000/pig_Output"
Counters:
Total records written : 0
Total bytes written : 0
Spillable Memory Manager spill count : 0
Total bags proactively spilled: 0
Total records proactively spilled: 0

Job DAG: job_1443519499159_0006

2015-10-05 13:06:06,192 [main] INFO  org.apache.pig.backend.hadoop.executionengine
.mapReduceLayer.MapReduceLau ncher - Success!

検証

以下に示すように、保存されたデータを確認できます。

ステップ1

まず、以下に示すように ls コマンドを使用して、 pig_output という名前のディレクトリ内のファイルをリストします。

hdfs dfs -ls 'hdfs://localhost:9000/pig_Output/'
Found 2 items
rw-r--r-   1 Hadoop supergroup          0 2015-10-05 13:03 hdfs://localhost:9000/pig_Output/_SUCCESS
rw-r--r-   1 Hadoop supergroup        224 2015-10-05 13:03 hdfs://localhost:9000/pig_Output/part-m-00000

*store* ステートメントの実行後に2つのファイルが作成されたことを確認できます。

ステップ2

*cat* コマンドを使用して、以下に示すように *part-m-00000* という名前のファイルの内容をリストします。

$ hdfs dfs -cat 'hdfs://localhost:9000/pig_Output/part-m-00000'
1,Rajiv,Reddy,9848022337,Hyderabad
2,siddarth,Battacharya,9848022338,Kolkata
3,Rajesh,Khanna,9848022339,Delhi
4,Preethi,Agarwal,9848022330,Pune
5,Trupthi,Mohanthy,9848022336,Bhuwaneshwar
6,Archana,Mishra,9848022335,Chennai

Apache-pig-storing-data

目次

Apache Pig-データの保存

構文

例

出力

検証

ステップ1

ステップ2