Our Community is getting an upgrade! To get everything ready for the relaunch, we’ll be placing the site in read-only mode starting September 21st.
We really appreciate your understanding while we get things set up behind the scenes. Catch up on all the exciting details about the move here.
Need help or have questions? Drop us a line at [email protected]
Created on
03-02-2023
03:13 AM
- last edited on
04-21-2026
12:42 AM
by
GrazittiAPI
On Cloudera Public Cloud the storage unit is GCS. Therefore, hive tables and any inserted data are stored on GCS rather than local disks ( like as on-prem ). But, this adds an overhead because when running a job, the node manager needs to get the data from gcs and cache it locally till the job is done. This eventually hits the performance on the env. Is there any way to move the data location from GCS to Local disks on Public Cloud clusters?
As written on the docs, datahub hdfs/disks spaces are temporary places but I would take this risk in favor of performance.
@steven-matison would really appreciate your help on this question. Thanks in advance.
Created 03-02-2023 06:44 AM
@fahed What you see with the CDP Public Cloud Data Hubs using GCS (or object store) is a modernization of the platform around object storage. This removes differences across aws, azure, and on-prem (when Ozone is used). It is a change by customer demand so that workloads are able to be built and deployed with minimal changes from on prem to cloud or cloud to cloud. Unfortunately that creates a difference you describe above, but those are risks we are willing to take ourselves in favor of modern data architecture.
If you are looking for performance, you should take a look at some of the newer options for databases: impala and kudu (this one uses local disk). Also we have Iceberg coming into this space too.
Created 03-02-2023 06:44 AM
@fahed What you see with the CDP Public Cloud Data Hubs using GCS (or object store) is a modernization of the platform around object storage. This removes differences across aws, azure, and on-prem (when Ozone is used). It is a change by customer demand so that workloads are able to be built and deployed with minimal changes from on prem to cloud or cloud to cloud. Unfortunately that creates a difference you describe above, but those are risks we are willing to take ourselves in favor of modern data architecture.
If you are looking for performance, you should take a look at some of the newer options for databases: impala and kudu (this one uses local disk). Also we have Iceberg coming into this space too.