Job Location: Bangalore/Bengaluru
Roles and Responsibilities
– Experience in building scalable/highly available distributed systems in production
– Expertise in Big Data Ecosystem with good experience in Java/Python/Scala, HDFS, Hive, Kudu, Spark, SQL and NoSQL databases (e.g. PostgreSQL, MS SQL, HBase, Cassandra, MongoDB)
– Expertise in MPP architecture and knowledge on MPP engine (Spark, Impala etc.)
– Understanding of stream processing with expert knowledge on Kafka and either Spark Streaming or Flink.
– Real-world experience with ETL/Data warehousing Solutions(NiFi, Streamsets, Talend, Informatica etc.)
– Knowledge of Cloud (Azure or AWS or Google Cloud) platform infrastructure and on premise to run and optimize distributed data applications
– Knowledge of Software Engineering best practices
– Identify right open source tools to deliver product features by performing research, POC/Pilot and/or interacting with various open source forums
– Engage with Product Management and Business to set your priorities and deliver awesome product features.
– Optimize solutions for performance and scalability.
– Load and Transform Data via using scripting languages and tools(e.g. Python, R, Java, Scala, Linux Shell, Sqoop)
– Collect and transform structured/unstructured data providing Self-serve capabilities for less technical analytics peers
– Identify gaps in the data processes and Drive improvement via continuous improvement loop (Data process Resilience)
– Designing, creating, testing and maintaining the complete data management and processing systems
– Discovering data acquisitions opportunities
– Improving data quality, reliability and efficiency of the individual components and the complete system
Submit CV To All Data Science Job Consultants Across India For Free

