An AI company is looking for a Platform Engineer. This is a remote and/or hybrid position.
Responsibilities:
- Cloud Infrastructure & Architecture
- Design and manage infrastructure across AWS, GCP, or Azure
- Build scalable systems for product services, AI workloads, and tooling
- Manage Kubernetes clusters, autoscaling, networking, and workloads
- Implement Infrastructure as Code using Terraform, Pulumi, or similar tools
- CI/CD & Developer Experience
- Build fast, secure deployment pipelines using GitHub Actions, GitLab CI, Jenkins, or ArgoCD
- Implement canary, blue-green, and rollback deployment strategiesCreate internal tools that improve engineering productivity
- Evaluate new technologies that improve platform speed and usability
- Observability & Reliability
- Build monitoring systems using Prometheus, Grafana, Datadog, or ELK
- Implement alerting, tracing, and centralized loggingTrack uptime, deployment frequency, latency, and platform health
- Improve debugging workflows across services and AI pipelines
- Performance & Scalability
- Design self-healing, fault-tolerant systems
- Optimize cost, throughput, and resource utilization
- Build autoscaling systems for changing AI and user demand
- Continuously remove bottlenecks across infrastructure layers
- Security & Compliance
- Enforce IAM, encryption, secrets management, and network security
- Manage patching, scanning, and hardening processes
- Support privacy-first infrastructure standards
- Collaborate with security teams to embed secure defaults
Requirements:
Experience:
- 3+ years in Platform Engineering, DevOps, or Site Reliability Engineering
- Experience operating production cloud-native systems at scale
- Strong Kubernetes and CI/CD experience
- Comfortable working cross-functionally in fast-moving teams
- Startup or growth-stage experience preferred
Technical Skills:
- Strong programming skills in Python, Go, TypeScript, or similar
- Deep knowledge of AWS, GCP, or Azure
- Hands-on Docker and Kubernetes expertise
- Infrastructure as Code experience (Terraform, Pulumi, CloudFormation)
- Monitoring and observability tooling experience
Soft Skills:
- Strong ownership mentality
- Excellent communication and documentation skills
- Systems thinking and architecture mindset
- Focus on developer experience and usability
Preferred Qualifications:
- AI / ML platform experience (GPU workloads, model serving, MLOps)
- Real-time systems or WebSocket infrastructure experience
- Developer portals such as Backstage or Port
- Chaos engineering and resilience testing knowledge
- Multi-cloud or hybrid-cloud experience
- SaaS, EdTech, or Career Tech background
Salary: $140,000 – $195,000
Hours: Full-Time – 40+
Zip Code: 10003
Job ID: #69736PE