Our Community is getting an upgrade! To get everything ready for the relaunch, we’ll be placing the site in read-only mode starting September 21st.
We really appreciate your understanding while we get things set up behind the scenes. Catch up on all the exciting details about the move here.
Need help or have questions? Drop us a line at [email protected]

Archives of Support Questions (Read Only)

This is an archived board for historical reference. Information and links may no longer be available or relevant
Announcements
This board is archived and read-only for historical reference. To ask a new question, please post a new topic on the appropriate active board.

Connect to CDH4.5 From Other Servers

avatar
Frequent Visitor

I have a CDH4.5 cluster, and I want to upload files into it from another server (e.g. database server).

 

With vanilla Hadoop and Hive, I can change the configuration files, pointing the namenode and metastore to remote services, and simply run:

 

dba@db-001$ hadoop fs -copyFromLocal /path/to/export.tsv
dba@db-001$ hive -e "load data local inpath '/path/to/export.tsv' into table test.my_table"

 

But how about CDH? What components should I install on other servers?

1 ACCEPTED SOLUTION

avatar
Guru

I think what you're describing is what we refer to as a "Gateway" machine.  On a cluster under Cloudera Manager's control, we allow you to add the "Gateway" role to a machine outside the cluster.  This installs the base CDH packages and deploys client configurations to that machine so that it can run regular hadoop commands like you describe and can upload files and run jobs against the cluster.

 

It sounds like your database server is already able to do this, can you clarify the question?

 

Regards.

View solution in original post

1 REPLY 1

avatar
Guru

I think what you're describing is what we refer to as a "Gateway" machine.  On a cluster under Cloudera Manager's control, we allow you to add the "Gateway" role to a machine outside the cluster.  This installs the base CDH packages and deploys client configurations to that machine so that it can run regular hadoop commands like you describe and can upload files and run jobs against the cluster.

 

It sounds like your database server is already able to do this, can you clarify the question?

 

Regards.