Scientific Dataset Engineering and ML Infrastructure
ZenteiQ
Description
Internship Opportunity
Scientific Dataset Engineering & ML Infrastructure
Location: Bengaluru, India
Duration: 3–6 months
Target: B.Tech, M.Tech, BS-MS, PhD students
What you will do :
- Implement geometry analysis, numerical computations, deduplication. Optimize for 500k+ cases.
- Quality Control & Validation: Design solution quality metrics, develop anomaly detection, build automated audit pipelines. Test edge cases thoroughly.
- Data Generation & Curation: Orchestrate parametric simulation variants, integrate with solvers, implement diversity-aware deduplication.
Requirements
Must Have:
Essential:
• B.Tech/M.Tech (final year) or PhD in Engineering, or Applied Math
• Strong numerical methods knowledge (FEM, FVM, FDM, convergence analysis)
• Hands-on coding: Python (NumPy, SciPy, Pandas), scientific computing experience
• Comfort with real solvers and scientific workflows (not just ML frameworks)
Highly Valued:
• C++ expertise (C++17+), memory optimization, performance tuning
• GPU computing (CUDA) or HPC experience (OpenMP, MPI)
• Hands-on with scientific solvers
• Published papers or significant open-source contributions
FAQ
Q: What domain is this?
A: Confidential. Disclosed after NDA signing.
Q: Do I need domain expertise?
A: Strong numerics and coding required. Domain knowledge is learnt.
Q: Mentorship level?
A: Weekly 1:1s, daily Slack, code reviews on every PR, open office collaboration.
Q: Remote work possible?
A: No.
Q: Will this lead to a full-time offer?
A: Full-time offers considered for proven track records in execution and code quality. Top performers (~15-20%) fast-tracked for R&D/SDE roles. Strong recommendations provided to all.
Benefits
Why This Internship Is Valuable
1.Access to Advanced Infrastructure & Research Support
2.Ownership & Visibility
3.Exceptional Mentorship & Learning Curve
4.Builds a profile aligned with industry research roles.
5.Real Research-Grade, Production-Impact Work
About ZenteiQ
-
