Job Location: Bengaluru
- Design, implement and maintain high performance big data infrastructure/systems & big data processing pipelines scaling to billions of structured and unstructured events daily
- Design, implement and maintain deep integration with up-stream systems
- Monitor performance of the data platform and optimize as needed
- Monitor and provide transparency into data quality across systems (accuracy, consistency, completeness, etc)
- Support product with the overall roadmap and ensure updates to senior leadership are 100% technically correct.
- Data analysis, understanding of business requirements and translation into logical pipelines & processes
- Conduct timely and effective research in response to specific requests (e.g. data collection, summarization, analysis, and synthesis of relevant data and information)
- Evaluate and prototype new technologies in the area of data processing
- Think quickly, communicate clearly and work collaboratively with product, data, engineering, QA and operations teams
- High energy level, strong team player and good work ethic
Technologies we use
- Google Cloud Platform (GCS, Dataflow, BigQuery, DataStudio)
- Python, R, Spark (PySpark)
- Kubernetes
- Airflow
- Git for source code management
- Jira
Qualifications
- BS in Computer Science or other technical discipline
- 5+ years of experience designing and developing big data processing systems using distributed computing
- Fluency in Python
- Expert knowledge in optimizing complex SQL queries
- Experience with job orchestration tools
- An affinity for automation
- Experience working with cloud platforms such as AWS & GCP
- Familiarity with networking and network application programming, including HTTP/HTTPS, JSON, and REST APIs
- Experience with at least one object oriented language (ex: Java)
- Strong OO design, data structure, and algorithm design skills
- Strong interest in emerging technologies: Spark, Hadoop, Hive, ElasticSearch, NoSQL
Nice to Have
Experience in the digital advertising and marketing industry
Experience in Scala and Spark
Experience with data science / modelling (TensorFlow and/or MXNet)
Familiarity with API development and design
Contribution to Open Source projects
Submit CV To All Data Science Job Consultants Across India For Free

