Job Location: Noida
Roles and Responsibilities
- Design and deliver scalable big data systems supporting both streaming and batch modes
- Implement and maintain the data-lake ecosystem to unify all diverse data sources.
- Build reliable data pipelines and ETLs to deliver data from diverse data sources.
- Work closely with stakeholders, including engineering, product, and analytics teams, to fulfill their requirements.
- Explore and assess novel technologies to address issues in the big data stack.
Desired Candidate Profile
- BS/MS or more in computer engineering/science or related experience
- 2+ years of industry experience in software development using Python, Java/Scala and SQL
- Understanding of distributed systems related to data processing and storage
- Knowledge of Cloud Technologies like GCP/S3, Glue, BigQuery, Athena
- Experience using data ETLs including Airflow or Nifi
- Familiar to Docker deployment, Kubernetes and CI/CD automation based on Gitflow
Preferred Qualifications
- Experience in Programming Data Pipelines via Python/Java and work with open source data pipelines tools like Airflow/Nifi/Spark/Glue
- Experience in working either AWS/Google Cloud Ecosystems
- Experience in data pipelines including Logstash or Filebeat or FluentD.
- Experience with stream processing such as Flink or Spark.
- Experience in NoSQL databases like Cassandra/Hbase/MongoDB and Columnstore DB and Row Store Databases like Sql Server etc is plus
Submit CV To All Data Science Job Consultants Across India For Free

