← Back to jobs
IIgnatiuz

AI Infrastructure and Platform Architect

Ignatiuz

Indore
Full-Time
1.2 - 1.8 LPA
3 experience

Description

Position Summary

We are seeking an experienced AI Infrastructure and Platform Architect to design, optimize, and manage scalable AI infrastructure and platforms across on‐premises, cloud, and hybrid environments.

The ideal candidate should have strong experience with GPU‐based systems, AI/ML platforms, infrastructure architecture, performance optimization, capacity planning, and production support. The role will work closely with AI/ML developers, DevOps engineers, data engineers, and solution architects to improve the performance, reliability, scalability, and cost efficiency of AI solutions.

Key Responsibilities

  • Design and manage AI infrastructure for model training, fine‐tuning, inference, computer vision, Generative AI, and LLM workloads.
  • Define CPU, GPU, RAM, VRAM, storage, networking, cooling, and power requirements.
  • Review existing hardware and platform performance and recommend upgrades or optimizations.
  • Perform capacity planning to support future workloads and minimize frequent hardware changes.
  • Build and maintain AI platforms using Linux, Docker, Kubernetes, GPU orchestration, and cloud services.
  • Configure and manage NVIDIA drivers, CUDA, cuDNN, TensorRT, and related AI acceleration technologies.
  • Monitor system health, GPU utilization, memory usage, storage performance, and network throughput.
  • Diagnose infrastructure failures, system crashes, performance bottlenecks, and platform outages.
  • Implement monitoring, alerting, backup, disaster recovery, security, and operational best practices.
  • Prepare architecture documents, hardware specifications, technical recommendations, and operational runbooks.
  • Support production deployment, troubleshooting, and continuous platform improvement.

AI Solution Optimization

The candidate should also be capable of:

  • Reviewing the end‐to‐end AI solution and identifying performance, architecture, and infrastructure gaps.
  • Recommending improvements to scalability, reliability, maintainability, and cost efficiency.
  • Supporting AI/ML developers with model training and experimentation environments.
  • Helping reduce training time through GPU optimization, distributed training, resource tuning, and efficient data pipelines.
  • Providing guidance on model accuracy, evaluation, hyperparameter tuning, and experimentation practices.
  • Improving model‐serving and inference performance.
  • Mentoring existing team members on AI infrastructure and production‐readiness best practices.

Eligibility Criteria

  • AI infrastructure and GPU-based computing
  • NVIDIA GPU architecture, CUDA, cuDNN, NCCL, and TensorRT
  • Linux administration
  • Docker and Kubernetes
  • PyTorch, TensorFlow, Hugging Face, or similar frameworks
  • Cloud and on-premises AI platforms
  • Infrastructure sizing and capacity planning
  • Performance monitoring and troubleshooting
  • High-performance storage and networking
  • MLOps, CI/CD, automation, and Infrastructure as Code
  • Monitoring tools such as Prometheus, Grafana, NVIDIA DCGM, or OpenTelemetry
  • AI solution architecture
  • Large Language Models and Generative AI
  • RAG and agentic AI systems
  • Distributed model training
  • Computer vision and edge AI
  • Model-serving platforms
  • AI performance benchmarking
  • FinOps and infrastructure cost optimization
  • High Performance Computing environments

About Ignatiuz

Ignatiuz is an Indian Microsoft Silver Partner technology consultancy offering SharePoint development, Office 365, Azure, and Google Cloud solutions.

Industry: TechnologyEmployees: 1000+Website