Use Apache Pig query results to enrich Hadoop Pig events.
On the system navigation tree, select System Properties.
Click Data Enrichment, then click Add.
On the Main tab, fill in the fields, then click the Source tab. In the Type field, select Hadoop Pig and fill in: Namenode host, Namenode port, Jobtracker host, and Jobtracker port.
Note
Jobtracker information is not required. If Jobtracker information is blank, NodeName host and port are used as the default.
On the Query tab, select the Basic mode and fill in the following information:
In Type, select text file and enter the file path in the Source field (for example,
/user/default/file.csv). Or, select Hive DB and enter an HCatalog table (for example,sample_07).In Columns, indicate how to enrich the column data.
For example, if the text file contains employee information with columns for SSN, name, gender, address, and phone number, enter the following text in the Columns field:
emp_Name:2, emp_phone:5. For Hive DB, use the column names in the HCatalog table.In Filter, you can use any Apache Pig built-in expression to filter data. See Apache Pig documentation.
If you defined column values above, you can group and aggregate that column data. Source and Column information is required. Other fields can be blank. Using aggregation functions require that you specify groups.
On the Query tab, select the Advanced mode and enter an Apache Pig script.
On the Scoring tab, set the score for each value returned from the single column query.
On the Destination tab, select the devices to which you want to apply enrichment.