Job Description
Are you ready to architect the digital landscape of 2026?
Nexus Future Systems is at the forefront of next-generation artificial intelligence. We are seeking a visionary AI Infrastructure Engineer to build the resilient, scalable, and high-performance computing environments that will power the future of generative AI. If you thrive on complexity and want to define the standards for AI operations in the coming decade, we want to hear from you.
In this pivotal role, you will bridge the gap between cutting-edge research and robust engineering, ensuring our AI models run at peak efficiency on massive distributed systems.
Responsibilities
- Architect Scalable AI Pipelines: Design and implement robust, cloud-native infrastructure to support large-scale machine learning model training and inference.
- Optimize Performance: Continuously monitor, analyze, and optimize system performance to reduce latency and increase throughput for AI workloads.
- Deploy MLOps Solutions: Establish CI/CD pipelines for machine learning models, ensuring automated testing, deployment, and rollback capabilities.
- Manage Resources: Utilize Kubernetes and containerization technologies to orchestrate microservices and manage resource allocation dynamically.
- Collaborate with Researchers: Partner with data scientists and AI researchers to translate theoretical models into production-ready infrastructure.
- Security & Compliance: Implement best practices for data security, privacy, and compliance in AI systems.
Qualifications
- Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field; PhD preferred.
- Experience: 5+ years of experience in software engineering, with at least 3 years specifically focused on machine learning infrastructure or MLOps.
- Programming: Strong proficiency in Python, Go, or Java, with deep knowledge of PyTorch and TensorFlow.
- Cloud Expertise: Proven experience deploying and managing workloads on AWS, Azure, or Google Cloud Platform (GCP).
- Containerization: Extensive experience with Docker and Kubernetes orchestration.
- Problem Solving: Demonstrated ability to troubleshoot complex distributed systems issues under pressure.