Our Community is getting an upgrade! To get everything ready for the relaunch, we’ll be placing the site in read-only mode starting September 21st.
We really appreciate your understanding while we get things set up behind the scenes. Catch up on all the exciting details about the move here.
Need help or have questions? Drop us a line at [email protected]

Community Articles

Find and share helpful community-sourced technical articles.
Announcements
Share your experience with Cloudera on G2 and get a $25 Amazon Gift card.
Hi, I'm CLEO! Something exciting is coming to the Community. Stay Tuned!
avatar
Rising Star

Setting Up a Data Science Platform on HDP using Anaconda

16600-dsc001-datascience-platform-on-hdp.png

Building a Data Science Platform using Anaconda needs to be able to

  • Launch PySpark jobs on the cluster
  • Synchronize python libraries from vetted public repositories
  • Isolate environments with specific dependencies to run production jobs using an older version of a package whilst simultaneously running new version of the package
  • Launching notebooks and PySpark jobs using different kernels such as Python_2.7, Python_3.x, R, Scala

Framework of the Data Science Platform

  • Private Repo Server
  • Edge Nodes
    • Dev
    • Test
    • Prod
    • Ansible
    • Git
    • Jenkins

Building blocks of the Data Science Platform

  • Anaconda
  • Ansible
  • Git
  • Jenkins
5,315 Views
Comments
avatar
New Member

And how to implement this, step how to install? how to install on existing HDP cluster?

Version history
Last update:
‎04-21-2026 06:29 AM
Updated by: