← Back to jobs
CChryselys

Sr. Associate - Data Engineering

Chryselys

Chryselys, Hitech City, Hyderabad, Telangana 500081
Full-Time
0.0 - 0.0 LPA
2-6 experience

Description

Data Engineering – Sr. Associate (Data Governance) ROLE OVERVIEW We are seeking a skilled Data Engineering Senior Associate with 4–6 years of hands–on experience to join our Life Sciences / Pharma data practice. In this role, you will design, build, and govern end–to–end data pipelines on Databricks, ensuring that data assets are accurate, compliant, and fit for consumption across regulated pharmaceutical environments. You will be a key contributor to our data governance framework, working closely with data stewards, compliance teams, and business stakeholders to implement robust standards for data quality, lineage, privacy, and lifecycle management.

KEY RESPONSIBILITIES

Data Ingestion & Pipeline Development

  • Design and implement scalable ingestion pipelines for both structured data (relational databases, flat files, EDI feeds) and unstructured data (clinical documents, lab reports, imaging metadata, free–text notes).
  • Build and maintain end–to–end ETL/ELT workflows on Databricks using PySpark, Spark SQL, and Delta Live Tables, covering ingestion, data quality checks, transformation, standardization, and serving layers.
  • Implement bronze–silver–gold (medallion) lakehouse architecture patterns with proper schema enforcement and data versioning.

Data Quality & Governance Implementation

  • Develop and enforce comprehensive DQ frameworks including completeness, accuracy, consistency, timeliness, and uniqueness checks using industry–wide used DQ tools
  • Implement DQ rules at multiple pipeline stages (ingestion, transformation, and aggregation) and surface metrics through automated dashboards and alerting.
  • Collaborate with data owners to define SLAs and remediation workflows for DQ failures.

Pharma Data – Privacy, Sensitivity, PHI & PII

  • Apply pharma-specific data privacy principles, ensuring proper handling of Protected Health Information (PHI) and Personally Identifiable Information (PII) in compliance with HIPAA, GDPR etc.
  • Implement data masking, tokenization, dynamic data redaction, and role-based access controls (RBAC) within Databricks Unity Catalog.
  • Work on classification of data sensitivity levels and enforce access policies across the data platform.

Data Lifecycle Management

  • Implement data versioning strategies using Delta Lake time-travel capabilities to support audit trails and point-in-time recovery.
  • Design and operationalize data retention and archival policies, aligning with regulatory retention schedules and storage cost optimization goals.
  • Establish and maintain end-to–end data lineage using Unity Catalog lineage features and integration with external cataloging tools (e.g., Apache Atlas, Alation, Collibra) – Good to have

REQUIRED QUALIFICATIONS
Education & Experience- Bachelor's or Master's degree in Computer Science, Information Systems, Data Engineering, or a related technical discipline.

  • 4–6 years of progressive experience in data engineering with a demonstrable focus on data governance in at least the last 2 years.

Eligibility Criteria

  • Bachelor's or Master's degree in Computer Science, Information Systems, Data Engineering, or a related technical discipline.
  • 4–6 years of progressive experience in data engineering with a demonstrable focus on data governance in at least the last 2 years.
  • Proficient in Databricks platform, including Unity Catalog, Delta Lake, Delta Live Tables, Databricks Workflows, and cluster management.
  • Strong Python (PySpark) and SQL
  • Proven experience designing end-to-end data pipelines from raw ingestion through to analytics-ready datasets.
  • Experience processing both structured (RDBMS, CSV, Parquet) and unstructured data (JSON, XML, PDFs, clinical notes, images/metadata).
  • Hands-on implementation of DQ frameworks; familiarity with any of the DQ tools is an added advantage.
  • Experience on AWS (S3, Glue, IAM) or Azure (ADLS Gen2, ADF, Entra ID); preferably both.
  • Hands-on experience with pharma data assets such as Medical affairs data, clinical trial data, EHR/EMR, lab data, adverse event reporting, or commercial data (sales, claims, payers, Open Data etc.).
  • Working knowledge of HIPAA, PHI/PII classification, and data privacy regulations applicable in the Life Sciences sector.
  • Experience implementing data versioning, time-travel queries, and audit trails.
  • Knowledge of data retention and archival strategies including tiered storage and automated lifecycle policies.
  • Exposure to data lineage tooling and metadata management platforms (Unity Catalog lineage, Apache Atlas, Alation, or Collibra) – Good to have

About Chryselys

Chryselys is an Indian healthtech company providing clinical research and data analytics solutions for pharmaceutical and life sciences industries.

Industry: Technology, HealthcareEmployees: 50+Website