Overview
About LockedIn AI
LockedIn AI is a real-time AI interview and meeting copilot trusted by more than one million users worldwide.
Our platform provides real-time, AI-powered assistance during live job interviews, coding assessments, and professional meetings, helping users communicate with greater confidence, clarity, and competence.
We are building an advanced AI-powered career platform and are looking for an experienced Platform Engineer to strengthen the infrastructure, internal tooling, and developer systems that support our engineering teams.
About the Role
As a Platform Engineer at LockedIn AI, you will design, build, and operate the internal platform used by our engineering, AI, data, and security teams.
You will own critical areas of our engineering infrastructure, including cloud architecture, Kubernetes, Infrastructure as Code, CI/CD pipelines, developer tooling, observability, security, automation, and production reliability.
This is a high-impact role. The platform and tools you build will directly influence how quickly engineers can develop, test, deploy, monitor, and scale services used by more than one million users.
The ideal candidate combines strong cloud infrastructure expertise with a product-focused approach to internal developer tooling. You view engineers as platform users and developer experience as a key measure of success.
Key Responsibilities
Cloud Infrastructure and Architecture
Design, build, and manage scalable cloud infrastructure using AWS, Google Cloud Platform, or Microsoft Azure.
Support compute, networking, databases, storage, IAM, serverless services, and AI model-serving infrastructure.
Architect reliable and cost-efficient infrastructure for product services, internal tools, and real-time AI workloads.
Build reproducible and auditable infrastructure using Terraform, Pulumi, CloudFormation, or Ansible.
Maintain infrastructure configurations through version control and automated deployment processes.
Improve cloud resource utilization through capacity planning, autoscaling, and cost optimization.
Kubernetes and Container Orchestration
Deploy, operate, maintain, and troubleshoot production Kubernetes clusters.
Package and manage services using Docker, Helm, and cloud-native technologies.
Implement autoscaling, service discovery, workload placement, and resource management.
Manage service mesh configurations and communication between microservices.
Support AI and machine learning workloads, including GPU infrastructure and model-serving endpoints.
Improve the reliability, security, and performance of containerized services.
CI/CD and Release Engineering
Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Argo CD, Jenkins, or similar platforms.
Automate application testing, validation, packaging, deployment, and rollback processes.
Implement canary releases, blue-green deployments, rolling deployments, and progressive delivery.
Create deployment safeguards that reduce production risk while maintaining engineering velocity.
Standardize deployment workflows across engineering, AI, and data teams.
Improve the speed and reliability of releases from code commit to production.
Internal Developer Platforms and Tooling
Build internal developer tools, APIs, command-line utilities, templates, and platform abstractions.
Reduce operational toil and eliminate repetitive engineering tasks through automation.
Improve local development, testing, deployment, debugging, and service onboarding workflows.
Build internal developer portals, service catalogs, and self-service infrastructure capabilities.
Evaluate and implement tools such as Backstage, Port, or custom developer experience platforms.
Gather developer feedback and continuously improve internal platform adoption and usability.
Observability, Monitoring and Alerting
Build monitoring and observability systems using metrics, logs, dashboards, and distributed traces.
Work with Prometheus, Grafana, Datadog, ELK, OpenSearch, or similar technologies.
Implement centralized logging and distributed tracing across microservices and AI pipelines.
Design intelligent alerts that reduce unnecessary noise and quickly surface production issues.
Monitor service availability, deployment frequency, cluster utilization, system performance, and developer experience metrics.
Create dashboards and operational runbooks that support rapid troubleshooting.
Reliability, Scalability and Performance
Design fault-tolerant and self-healing platform architecture.
Improve service availability during traffic spikes, infrastructure failures, and partial outages.
Identify and resolve latency, throughput, capacity, and resource bottlenecks.
Build autoscaling and capacity management systems for variable user and AI inference workloads.
Support incident response, root cause analysis, and long-term reliability improvements.
Implement platform standards that improve production resilience across engineering teams.
Security and Compliance
Implement least-privilege IAM policies, network segmentation, encryption, and secure access controls.
Manage secrets, credentials, audit logs, vulnerability scanning, and infrastructure patching.
Embed security checks into CI/CD pipelines and developer workflows.
Collaborate with security teams to strengthen cloud and application infrastructure.
Protect user data and support LockedIn AI’s privacy-first product standards.
Maintain secure configurations for Kubernetes, cloud services, databases, and internal tools.
Collaboration and Documentation
Partner with engineering, AI/ML, data, product, and security teams.
Understand team requirements and build platform capabilities that improve productivity.
Document platform architecture, operational procedures, runbooks, and infrastructure decisions.
Create architectural decision records and technical implementation guides.
Promote platform engineering best practices and a platform-as-a-product culture.
Stay current with cloud-native technologies, developer experience tools, and infrastructure practices.
Required Qualifications
Three or more years of experience in Platform Engineering, DevOps, Site Reliability Engineering, Cloud Engineering, or Production Infrastructure.
Experience building and operating cloud-native infrastructure in production environments.
Strong knowledge of AWS, GCP, or Azure.
Hands-on experience managing Docker and Kubernetes environments.
Experience using Terraform, Pulumi, CloudFormation, Ansible, or similar Infrastructure as Code tools.
Experience building CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or Argo CD.
Proficiency in Python, Go, TypeScript, or another programming language used for automation and infrastructure tooling.
Experience with Helm, service meshes, autoscaling, and Kubernetes resource management.
Knowledge of Prometheus, Grafana, Datadog, ELK, OpenSearch, or similar observability tools.
Understanding of distributed systems, cloud networking, IAM, security, and production reliability.
Experience collaborating with engineering, data, AI, and security teams.
Strong written and verbal communication skills.
A formal degree is not required. Demonstrated platform engineering expertise, systems thinking, and production experience are valued more than credentials.
Preferred Qualifications
Experience building infrastructure for AI or machine learning products.
Experience supporting GPU workloads, model registries, inference endpoints, and MLOps pipelines.
Knowledge of real-time systems, WebSockets, streaming audio, or latency-sensitive applications.
Experience building internal developer portals or service catalogs using Backstage, Port, or custom solutions.
Familiarity with GitOps and progressive delivery workflows.
Experience with chaos engineering and proactive resilience testing.
Experience managing multi-cloud or hybrid-cloud environments.
Background in SaaS, career technology, education technology, or consumer AI products.
Open-source contributions related to infrastructure, Kubernetes, DevOps, or platform engineering.
Previous startup or early-stage company experience.
Core Skills and Keywords
Platform Engineering, Cloud Infrastructure, AWS, GCP, Azure, Kubernetes, Docker, Terraform, Pulumi, Infrastructure as Code, CI/CD, GitHub Actions, GitLab CI, Jenkins, Argo CD, GitOps, Helm, Service Mesh, Prometheus, Grafana, Datadog, ELK, OpenSearch, Python, Go, TypeScript, Developer Experience, Internal Developer Platform, DevOps, Site Reliability Engineering, Cloud Security, Observability, Distributed Systems, Microservices, Automation, High Availability, Scalability and AI Infrastructure.
What We Offer
Competitive compensation of $140,000–$195,000 USD per year.
Meaningful early-stage equity.
Remote-first work for candidates based in the United States.
Optional hybrid and co-working opportunities in New York City.
Direct ownership of cloud infrastructure and internal developer platforms.
The opportunity to support an AI platform used by more than one million users.
A fast-moving environment with significant technical autonomy.
Close collaboration with engineering, AI, data, security, and leadership teams.
Why Join LockedIn AI?
You will build the engineering foundation that every product team and AI system depends on.
Your work will improve developer productivity, deployment speed, platform reliability, infrastructure security, and the performance of services supporting more than one million users.
This role provides the opportunity to work at the intersection of platform engineering, cloud infrastructure, artificial intelligence, developer experience, DevOps, and real-time systems.
How to Apply
Submit your application through the LockedIn AI careers page:
Apply for the Platform Engineer position here:
https://www.lockedinai.com/careers/platform-engineer
Please include:
Your resume or CV
A brief explanation of why you want to join LockedIn AI
Whether you have tried the LockedIn AI application
Your thoughts on what could be improved
Relevant GitHub repositories, portfolio projects, open-source contributions, or technical writing
Equal Opportunity Employment
LockedIn AI is committed to creating a diverse and inclusive workplace. We welcome candidates from all backgrounds, identities, and professional experiences.
Employment decisions are based on qualifications, experience, demonstrated skills, merit, and business requirements.