Big Data Project

€ 250.00 — 750.00 EUR
Need: We are looking to form a team of two skilled developers that are willing to work on a big data project for the three different tasks that we will explain in a bit.

Objective: The objective is to scale the platform for future needs as well as address limitations of retrieving data by building a mini search engine with Elasticsearch.
Minimizing the cost is important factor which motivated the use of open-source tools such as: Apache NiFi
Apache Spark
Ellasandra (Cassandra + Elastic search)

User stories:
User story 1: As a user, I want to be able to import a heavy csv file to Cassandra.

User story 2: As a user, I want to be able to search for a specific field (mail, id ...) and get the ten first occurrences displayed.

User story 3: As a user I want to be able to export a csv file that has the result of the query I wrote in the search engine.

Explanation of the role of the three project main functionalities:
The user will have access to a simple graphical user interface that will let him choose between:
- Import: has as objective to allow importing heavy datafiles
- Export: has as objective to generate a text file with respect to some filter (could be a table name, or a property) specified and fetched from the search input field
- Search: has as objective to filter the data and render at most 10 data rows that matches the search query and render it to the user.

Note: This project generates only one view for the user based on the input in the search field.

Expected features:
1) Simple graphic interface that contains the 3 main Sub functionalities of this interface:
button import, button export (export can be a collection, depends on the filter and the query), Search field (to filter).
• Level of priority: Medium
• Expected programming skills: HTML / CSS / Flask(preferably but not a must)
2) Settle Elassandra cluster (2 databases).
Load huge data file into the cluster in order to achieve distributed storage.
And finally export the result of a specific query to a csv file.
• Level of priority: High
• Expected programming skills: python, Elasandra NoSQL (Cassandra + Elasticsearch), Apache NiFi

3)Server side for data processing: This will have two main and separate goals:
First: If one of the special fields in already existent, upsert into already existing record the missing fields from the new record and vice-versa.

Second: Implement The search algorithm that will get the expected row, table or even field
• Level of priority: high
• Expected programming skills: python, Apache Spark, Elasandra (Cassandra+Elasticsearch)
Expected result: The project will be deemed successful if we see that the user stories are met, and the database fields are being updated as explained in the section before.

Similar Freelance jobs:

Problem With Spark Ar Filter
Need filter for instagram to be fixed. The image should be inside but instead it is cut off as seen in the image
Full Description of problem with Spark AR filter
Elasticsearch Query
Need to get count of documents based condition We do have an index having customers and orders information , we need to retrieve count of unique customers which have placed on order on particular range and unique 1. Get unique count of customers which have ordered today itself - it should not have any orders placed earlier. 2. Get unique count of customers which have not orders in particualr range but placed order at particualr date - festive date 3.…
Full Description of ELasticsearch Query
Sr Graph Db Expert W/ Aws Neptune Preferred + Nodejs
I'm looking for a Senior Graph DB expert with major experience with Gremlin, configuring and working through scaling issues with Neptune. Also, significant experience with NodeJS in data extracting from the websites.
Full Description of Sr Graph DB Expert w/ AWS Neptune Preferred…
Spark Scla
Expertise in designing and deployment of Hadoop Cluster and different analytical tools including Pig, Hive, HBase, Sqoop, Kafka Spark with Cloudera distribution. Working on a live 20 nodes Hadoop cluster running on CDH4.4. Working with highly unstructured and semi structured data of 40 TB in size (120 TB with replication factor of 3) Managing external tables in Hive for optimized performance. Very good understanding of Partitions and Bucketing in Hive Developed Spark scripts using Scala as per the requirement using…
Full Description of Spark Scla
Simple Assignment Using Hive And Pig
All details are mentioned in the attached doc. Please check before bidding. It will take you 4 to 5 hour to complete it
Full Description of Simple assignment using hive and pig
Distributed And Scalable Big Data Project (apache Spark ,apache Nifi, Cassandra, Elasticsearch)
We look for two big data developers to work on developing: - distributed Cassandra (huge data entry) , process data with spark(enhance parallel computing), a search engine( Elasticsearch) - simple graphic user interface.
Full Description of Distributed and scalable Big Data Project (Apache Spark…
Elasticsearch Consulting -- 2
I am looking for an Elasticsearch Expert having many years industry work experience for a long-term consulting. Our Elasticsearch System requires the following functionality, 1. Heavy search query from client side, i.e full-text search, keyword search, Regex search etc. 2. The search latency should not be more than 1 sec for full-text and keyword search, but for regex it can be a bit slow but not more than 3-4 sec. 3. We will ingest data to the server 24/7 i.e…
Full Description of Elasticsearch Consulting -- 2
Python Engineer
Experience in Python is a must. Experience on C# would be an added advantage. ● Proficiency with fundamental front-end languages such as HTML, CSS, and JavaScript, Jquery. ● Familiarity with JavaScript frameworks such as Vuejs, Angular JS, React. ● Beneficial to have working experience on POS Systems or any other hardware embedded & IoT systems. ● Worked on enterprise grade systems. ● Have designed web services (RESTful web services). ● Know how to scale systems that have database bottlenecks etc.…
Full Description of Python Engineer
Mongodb Big Data Python
i have a database on mongodb, and i have multiple tables ,, sometimes i have tables that have points in common, and so i want the table fields to be updated from their common data and which are not filled in their fields
Full Description of Mongodb big data python
© 2006 — 2026 hirelancer.com is an affiliate website, listing the freelance projects. We are collaborating with other sites like getFreelancer.com. Feel free to bid on any project and good luck! You will need register first, before bidding or posting projects. Basic memberhsip is always free. We might earn commissions from referring you. You will never ever pay additional fees for we referred you.