- Scala -
2.12.8 - SBT -
1.2.8 - Spark -
spark-2.4.3-bin-without-hadoop-scala-2.12 - Hadoop -
3.2.0 - Kafka -
2.2.1 - Docker -
18.09.5
- Install Apache Kafka
- I preffer to run it on separate VM;
- Download and setup Apache Hadoop 3.2.0;
- Download and setup the Apache Spark binaries;
- Format HDFS partition:
./bin/hdfs namenode -format
- Run DFS (Name and Data nodes of HDFS):
- Assuming you're in Hadoop HOME:
./bin/start-dfs.sh;
- Assuming you're in Hadoop HOME:
- Get the project sources:
git clone https://github.com/idenisovs/bigdata-challenge.git; - Assembly the project sources:
sbt assembly; - Build the Docker image for producer:
docker build --tag device-sim .- See the
Dockerfilesfor details;
- See the
- Run N-th number of producers (assuming the Kafka is up and running):
docker-compose up --scale device-sim=3- You can run the single producer on host machine:
./run-producer.sh
- Run consumer:
./run-cosumer.sh
- Observe the content of HDFS here: http://localhost:9870/explorer.html
root
|
+---> core (shared objects)
|
+---> device-simulator (Producer)
|
+---> process-job (Consumer)
messages
device sim 1 ----------> +-------+
| |
device sim 2 ----------> | Kafka |
| |
device sim 3 ----------> +-------+
|
| messages
V
+---------------+
| Spark |
| +-----------+ |
| |process-job| |
| +-----------+ |
+---------------+



