Our Community is getting an upgrade! To get everything ready for the relaunch, we’ll be placing the site in read-only mode starting September 21st.
We really appreciate your understanding while we get things set up behind the scenes. Catch up on all the exciting details about the move here.
Need help or have questions? Drop us a line at [email protected]
Created 10-18-2016 08:38 AM
i'm not able to create data frames with the information provided in the below link http://spark.apache.org/docs/latest/sql-programming-guide.html#data-sources I'm using HDP2.3-Pig & Hive Rev6 zip latest version may i know what is the other way to create data frames checked different forums but unable to get it done need your support to come over the issue.
Created 10-18-2016 01:29 PM
I assume you are running scala/spark code? Also, just as an FYI, the link that you shared above is for the latest version of Spark (which is currently 2.0.1). HDP 2.3 is an older version of the Hortonworks platform and it runs Spark 1.4.1. If you want to use the latest version of Spark, you can upgrade to HDP 2.5.
But you should not need to upgrade to create a dataframe within Spark. Here's spark/scala code that I used within HDP 2.3 to generate a sample dataframe:
val df = sqlContext.createDataFrame(Seq(
("1111", 10000,"M"),
("2222", 20000,"F"),
("3333", 30000,"M"),
("4444", 40000,"F")
)).toDF("id", "income","gender")
df.show()Try this out and let me know if it works.
Created 10-18-2016 01:29 PM
I assume you are running scala/spark code? Also, just as an FYI, the link that you shared above is for the latest version of Spark (which is currently 2.0.1). HDP 2.3 is an older version of the Hortonworks platform and it runs Spark 1.4.1. If you want to use the latest version of Spark, you can upgrade to HDP 2.5.
But you should not need to upgrade to create a dataframe within Spark. Here's spark/scala code that I used within HDP 2.3 to generate a sample dataframe:
val df = sqlContext.createDataFrame(Seq(
("1111", 10000,"M"),
("2222", 20000,"F"),
("3333", 30000,"M"),
("4444", 40000,"F")
)).toDF("id", "income","gender")
df.show()Try this out and let me know if it works.
Created 10-18-2016 02:03 PM
Thanks for the above it worked for me I was able to create the dataframe. My approach to access .json file is wrong or any thing that I need to try with.