The new docs.trellix.com features a modernized UI and AI-powered conversational search. Content is currently available in English, with additional languages launching in early November 2026. We hope you enjoy the updated experience.

Enrich events with Hadoop Pig

Prev Next

Use Apache Pig query results to enrich Hadoop Pig events.

  1. On the system navigation tree, select System Properties.

  2. Click Data Enrichment, then click Add.

  3. On the Main tab, fill in the fields, then click the Source tab. In the Type field, select Hadoop Pig and fill in: Namenode host, Namenode port, Jobtracker host, and Jobtracker port.

    Note

    Jobtracker information is not required. If Jobtracker information is blank, NodeName host and port are used as the default.

  4. On the Query tab, select the Basic mode and fill in the following information:

    1. In Type, select text file and enter the file path in the Source field (for example, /user/default/file.csv). Or, select Hive DB and enter an HCatalog table (for example, sample_07).

    2. In Columns, indicate how to enrich the column data.

      For example, if the text file contains employee information with columns for SSN, name, gender, address, and phone number, enter the following text in the Columns field: emp_Name:2, emp_phone:5. For Hive DB, use the column names in the HCatalog table.

    3. In Filter, you can use any Apache Pig built-in expression to filter data. See Apache Pig documentation.

    4. If you defined column values above, you can group and aggregate that column data. Source and Column information is required. Other fields can be blank. Using aggregation functions require that you specify groups.

  5. On the Query tab, select the Advanced mode and enter an Apache Pig script.

  6. On the Scoring tab, set the score for each value returned from the single column query.

  7. On the Destination tab, select the devices to which you want to apply enrichment.