Our Community is getting an upgrade! To get everything ready for the relaunch, we’ll be placing the site in read-only mode starting September 21st.
We really appreciate your understanding while we get things set up behind the scenes. Catch up on all the exciting details about the move here.
Need help or have questions? Drop us a line at [email protected]
Created on 12-09-2018 11:26 PM - edited 09-16-2022 06:58 AM
Hi All, I would like to implement a real time data feed between a webserver and hadoop server. I plan to use flume to read the web log files real time and target is hdfs/Hive,
Questions are:
1. I need a checklist of what to prepare for the security like, firewalls etc.
2. Are there any hadoop agent I need to install in the webser server
3. Once data is available now in hive, I will have a regular job to process the data using Impala then once processed I will have a list of suggestions/messages for a particular web user. How do I send the info back to that specific web users web page?
Thank you
Created 04-09-2019 06:53 AM
Created 04-09-2019 06:53 AM
Created 02-27-2020 11:43 PM
Thanks.
Implemented Flume to fetch the weblogs that was transferred to the Hadoop edge server upto HDFS.
Also, due to firewall challenges and security implementations and lack of test environment, used an alternative solution of using Zena job scheduler to transfer the log files from ATM machines and mobile web app logs to hadoop edge server.
Kafka came as a big challenge since we are using LDAP thus security and authentication issues quickly cropped up.
Kudus to your suggestion!