← Back to jobs
GGlobal Data Transformation Company

Associate Data Engineer

Global Data Transformation Company

Lower Parel, Mumbai
Full-Time
0-6months experience

Description

Location: Mumbai Experience: 0-6months Technologies / Skills: Advanced SQL, Python and associated libraries like Pandas, Numpy etc., Pyspark , Shell scripting, Data Modelling, Big data, Hadoop, Hive, ETL pipelines.

Responsibilities: • Proven success in communicating with users, other technical teams, and senior management to collect requirements, describe data modeling decisions and develop data engineering strategy. • Ability to work with business owners to define key business requirements and convert to user stories with required technical specifications. • Communicate results and business impacts of insight initiatives to key stakeholders to collaboratively solve business problems. • Working closely with the overall Enterprise Data Analytics Architect and Engineering practice leads to ensure adherence with the best practices and design principles. • Assures quality, security and compliance requirements are met for supported area. • Design and create fault-tolerance data pipelines running on cluster • Excellent communication skills with the ability to influence client business and IT teams • Should have design data engineering solutions end to end. Ability to come up with scalable and modular solutions

Required Qualification: • 0-6months of hands-on experience Designing and developing Data Pipelines for Data Ingestion or Transformation using Python (PySpark)/Spark SQL in AWS cloud • Experience in design and development of data pipelines and processing of data at scale. • Advanced experience in writing and optimizing efficient SQL queries with Python and Hive handling Large Data Sets in Big-Data Environments • Experience in debugging, tunning and optimizing PySpark data pipelines • Should have implemented concepts and have good knowledge of Pyspark data frames, joins, caching, memory management, partitioning, parallelism etc. • Understanding of Spark UI, Event Timelines, DAG, Spark config parameters, in order to tune the long running data pipelines. • Experience working in Agile implementations • Experience with building data pipelinesin streaming and batch mode. • Experience with Git and CI/CD pipelines to deploy cloud applications • Good knowledge of designing Hive tables with partitioning for performance.

Desired Qualification: • Experience in data modelling • Hands on creating workflows on any Scheduling Tool like Autosys, CA Workload Automation • Proficiency in using SDKsfor interacting with native AWS services • Strong understanding of concepts of ETL, ELT and data modeling.

Eligibility Criteria

0-6months

About Global Data Transformation Company

Global Data Transformation Company is an Indian data engineering firm building infrastructure for data extraction, transformation, and loading using SQL and AWS.

Industry: TechnologyEmployees: 1000+Website