Created 09-12-2017 07:44 PM
Partitioner is not invoked when used in oozie mapreduce action (Creating workflow using HUE). But works as expected when running using hadoop jar commad in CLI,
I have implemented secondary sort in mapreduce and trying to execute it using Oozie (From Hue).
Though I have set the partitioner class in the properties, the partitioner is not being executed. So, I'm not getting output as expected.
The same code runs fine when run using hadoop command.
And here is my workflow.xml
<workflow-app name="MyTriplets" xmlns="uri:oozie:workflow:0.5"> <start to="mapreduce-598d"/> <kill name="Kill"> <message>Action failed, error message[${wf:errorMessage(wf:lastErrorNode())}]</message> </kill> <action name="mapreduce-598d"> <map-reduce> <job-tracker>${jobTracker}</job-tracker> <name-node>${nameNode}</name-node> <configuration> <property> <name>mapred.output.dir</name> <value>/test_1109_3</value> </property> <property> <name>mapred.input.dir</name> <value>/apps/hive/warehouse/7360_0609_rx/day=06-09-2017/hour=13/quarter=2/,/apps/hive/warehouse/7360_0609_tx/day=06-09-2017/hour=13/quarter=2/,/apps/hive/warehouse/7360_0509_util/day=05-09-2017/hour=16/quarter=1/</value> </property> <property> <name>mapred.input.format.class</name> <value>org.apache.hadoop.hive.ql.io.RCFileInputFormat</value> </property> <property> <name>mapred.mapper.class</name> <value>PonRankMapper</value> </property> <property> <name>mapred.reducer.class</name> <value>PonRankReducer</value> </property> <property> <name>mapred.output.value.comparator.class</name> <value>PonRankGroupingComparator</value> </property> <property> <name>mapred.mapoutput.key.class</name> <value>PonRankPair</value> </property> <property> <name>mapred.mapoutput.value.class</name> <value>org.apache.hadoop.io.Text</value> </property> <property> <name>mapred.reduce.output.key.class</name> <value>org.apache.hadoop.io.NullWritable</value> </property> <property> <name>mapred.reduce.output.value.class</name> <value>org.apache.hadoop.io.Text</value> </property> <property> <name>mapred.reduce.tasks</name> <value>1</value> </property> <property> <name>mapred.partitioner.class</name> <value>PonRankPartitioner</value> </property> <property> <name>mapred.mapper.new-api</name> <value>False</value> </property> </configuration> </map-reduce> <ok to="End"/> <error to="Kill"/> </action> <end name="End"/>
When running using hadoop jar command, I set the partitioner class using JobConf.setPartitionerClass API.
Not sure why my partitioner is not executed when running using Oozie. Inspite of adding
<property> <name>mapred.partitioner.class</name> <value>PonRankPartitioner</value> </property>
Created 09-12-2017 07:46 PM
Created 09-16-2017 03:08 AM
I am also searching for the same fix , let me know if you have found one .
Created 09-16-2017 03:19 AM
Solved this issue by re-writing the mapreduce job using new API's.
The property used in oozie workflow for partitioner was mapreduce.partitioner.class.