
Data Engineer - Python
Global Data Transformation Company
Description
We are looking for a Python-centric Data Engineer who can design and maintain scalable, high-performance data pipelines for large and complex datasets. The ideal candidate brings strong object-oriented Python programming skills, experience with distributed data systems (like Hadoop or Spark) and a mindset for building modular, reusable, and testable data solutions. While exposure to cloud technologies (AWS preferred) is valuable, this role emphasizes deep Python-based data engineering over platform administration. Design, develop, and optimize data ingestion and transformation pipelines using Python. Apply OOP principles to build modular, maintainable, and reusable data components. Work with distributed systems (e.g., Hadoop/Spark) for processing large-scale datasets. Develop and enforce data quality, testing, and validation frameworks. Collaborate with analysts, data scientists, and product teams to ensure data availability and reliability. Participate in code reviews, CI/CD workflows, and infrastructure automation to maintain high engineering standards. Contribute to ongoing evolution of data architecture and help integrate with cloud-based data ecosystems.
Eligibility Criteria
Bachelor's or master's degree in engineering or technology or related field. 3–6 years of hands-on experience in data engineering or backend development. Proven track record of Python-based pipeline development and distributed data processing. Strong foundation in data modeling, data quality, and pipeline orchestration concepts. Excellent problem-solving and communication skills, with an ownership-driven mindset.
About Global Data Transformation Company
Global Data Transformation Company is an Indian data engineering firm building infrastructure for data extraction, transformation, and loading using SQL and AWS.
